International Journal of Communication 20(2026) Rethinking Sociotechnical Harms 

 

Rethinking Sociotechnical Harms of Large Foundation Models: A Critical Reflection Through Scale Hacking

 

ANGEL HWANG

University of Southern California, USA

 

With the rise of large general-purpose models, understanding the societal impact of AI systems becomes more difficult yet critical. Many AI harms are uniquely tied to scale, where certain sociotechnical risks only emerge or become intensified when widely adopted by users. Assessing these risks before large-scale model deployment is desirable; in practice, it is extremely challenging to study such harms without involving a large number of users. This paper introduces scale hacking—a set of methodological strategies from the human-computer interaction (HCI) literature designed to extend research insights without requiring massive sample sizes—as a lens to rethinking such harms. I reflect on how we applied these scale hacking techniques to examine the impact of foundation models in an applied setting (i.e., freelance economy), highlighting three effective approaches: (1) triangulating across multiple methods and data streams, (2) including participants with varying degrees of experience over both short and long time frames, and (3) focusing not on specific AI applications but on user practices such as disclosure of AI use.

 

Keywords: sociotechnical harm, large foundation model, general-purpose model, AI harm, scale hacking

 

 

As large foundation AI models become increasingly embedded in digital infrastructures and applications, scholars across disciplines have raised concerns about their sociotechnical harms (Rauh et al., 2022). Unlike conventional machine learning systems designed for narrowly defined tasks, foundation models, such as GPT, Claude, and Gemini, are “trained on broad data that can be adapted to a wide range of downstream tasks” (Bommasani et al., 2022, p. 1). While this general-purpose approach enables robust performance across multiple domains, it also complicates efforts to anticipate the novel forms of harm these models may produce (Bommasani et al., 2022; Shelby et al., 2023; Zhou et al., 2024).

 

Existing research highlights that many of these harms of large foundation models are inherently tied to scale (Domínguez Hernández et al., 2024; Shelby et al., 2023; Weidinger et al., 2022). These systems are precisely designed to be at their massive scales in order to support broad adoption, while expansion in use amplifies both existing and novel risks. Ideally, practitioners would assess these risks before models become widely accessible. In practice, it is extremely difficult to anticipate such harms in advance, as there is limited empirical precedent to inform realistic modeling of large-scale, sociotechnical interactions. Because these systems operate at scales that intersect multiple infrastructures—economic, informational, social, and more—their downstream effects are nearly impossible to predict in empirical, controlled settings.

 

This paper aims to contribute to addressing this challenge by introducing the concept of scale hacking, a line of thinking and techniques from the human–computer interaction (HCI) literature that explore how the potential broad-reaching impacts of technical systems can be studied without requiring access to a massive pool of testing subjects. We apply scale hacking to design two complementary studies that investigate the impact of large foundation models in an applied setting (i.e., a freelance labor market) and critically reflect on the extent to which scale hacking helps researchers meaningfully examine societal impacts of foundation models, particularly in contexts where real-world deployment has yet to reach full scale.

 

This paper makes three primary contributions. First, by introducing scale hacking to the audience of the International Journal of Communication, we aim to help researchers studying the societal impact of foundation models consider methodologies and frameworks beyond those traditionally employed within their fields. Second, we show how scale hacking can be operationalized in empirical research, prompting communication scholars to experiment with their own forms of scale hacking by applying this concept to their research methods. Finally, by foregrounding techniques from HCI—where design implications of technological systems are often central to research contributions—we seek to promote actionable, design-oriented thinking within communication research, informing the development and governance of future large-scale AI systems.

 

Conceptualizing “Scale”

 

In this paper, scale is defined in terms of technological adoption—specifically, the extent to which users incorporate AI systems into their everyday work and life (Larsen, 2021). We deliberately keep the definition of scale simple, enabling us to better interrogate the social and institutional consequences of widespread AI adoption.

 

Scale Hacking in Human-Computer Interaction (HCI) Research

 

Although the term “scale hacking” was formally introduced by Brown et al. (2017), scale has long been central to sociotechnical research. For instance, studying social media’s impact requires an understanding of the broader dynamics of online social networks and communities. To move beyond studying user-technology interaction in an isolated fashion, research should consider sociotechnical influences across various settings, engaging with the multifaceted aspects of infrastructures and artifact ecologies simultaneously (Korsgaard et al., 2022). Hence, scaling in HCI research involves not only the number of users but also the diversity of contexts and the multiplicity of interconnected technologies (Bødker & Klokmose, 2012; Brown et al., 2017). In this sense, scale hacking is about ensuring that design knowledge transcends individual use cases to inform interactions in more complex environments (Moradi et al., 2018).

 

Prior research outlines four lines of methodological innovation that support scaling HCI knowledge (Brown et al., 2017; Kraut et al., 1996; Notess & Blevis, 2004; Wiberg & Stolterman, 2021): (a) considering the role of culture in design, (b) shifting from discrete interaction techniques to broader gestalts of interaction, (c) leveraging data in the design process, and (d) employing longitudinal studies to observe technology-in-use over time. These approaches enable researchers to generate robust insights even without access to large-scale empirical data, highlighting that scale hacking is less about sample size and more about methodological strategies that extend the relevance of findings beyond their original context. In contrast, traditional user-centered methods focus on individuals or small groups, which can fall short in capturing the dynamics as large-scale AI systems operating across distributed, data-rich ecosystems (Xu et al., 2023). Addressing challenges posed by such systems requires rethinking HCI frameworks to better accommodate their emergent, sociotechnical complexities (Ozmen Garibay et al., 2023).

 

Case Study: Impact of AI Adoption on Freelance Economies

 

We applied scale hacking techniques in two studies examining the impact of adopting AI-powered tools on two freelance platforms (i.e., Bēhance and Upwork), where the extent to which freelancers applied AI for work varies. We conducted two series of mixed-methods studies, combining large-scale qualitative analyses with quantitative methods, to study small groups of select participants. Our findings illustrate that as AI adoption scales, it creates new norms on such platforms. These norms benefit certain professionals (e.g., tech experts are compensated more when they claim to use AI for work), but cause harm to others (e.g., creative professionals are paid less when they adopt AI for content production). As general-purpose models become more integrated into freelance workflows, they reshape the perceived value of work within online labor markets.

 

Investigating Harms of Large Foundation Models Through “Scale Hacks”

 

We applied four common scale hacking techniques to our study. We examine how these techniques enabled us to explore the downstream effects of applying large AI systems in the freelance landscape.

 

Triangulating Methods to Go Beyond Individual Design

 

Rather than focusing on individual user experiences, researchers can adopt methods—and ideally, triangulate insights from multiple methods—that provoke reflection from multi-stakeholders. Prior work has demonstrated that applying and combining methods, such as design fictions (Coulton et al., 2017), speculative research (Sengers et al., 2021), and other critical approaches (Farias et al., 2022; Wong & Nguyen, 2021) can reveal latent concerns, moral dilemmas, and power asymmetries that are not immediately visible in individuals’ interaction with technology. Motivated by this principle, we employed a mixed-methods study design, combining field data collection (e.g., analyzing large-scale datasets retrieved from platform APIs), controlled behavioral experiments, and critical methods (including interviews and workshops) that encourage participants to reflect on near- and long-term implications of AI adoption. This triangulation revealed sociotechnical harms not only at the individual level but across broader systemic patterns in the online labor market. 

 

Data as Design Material

 

Rather than treating data as a static end product, researchers could see it as a design material to shape, contrast, and interrogate (Alfaras et al., 2020; Benjamin et al., 2021). In our study, unresolved questions emerging from our qualitative interviews informed the design of our large-scale field data collection. Each data stream offered unexpected insights, revealing nuances that went beyond our initial research plan. By comparing creators and non-creative freelancers, we identified divergent patterns in how AI tools were perceived and valued. These insights did not result from the sheer volume of data but from sampling, segmenting, and triangulating across complementary datasets.

 

Temporal Scaling: Envisioning Technological Impact Over Time

 

A key goal of scale hacking is to anticipate how technology influences may evolve over time and across contexts (Brown et al., 2017; Wiberg & Stolterman, 2021). This temporal dimension is particularly critical in the context of AI, as many of its potential harms emerge or intensify over time, shaped by shifting social norms, expectations, and market conditions (Lee & Whitley, 2002; Shelby et al., 2023; Weidinger et al., 2022). Hence, our study spans multiple years to address this temporal dimension, beginning in early 2023 during the initial wave of widespread public enthusiasm for generative AI tools. Furthermore, we studied professionals at various career stages and with differing levels of AI adoption and platform engagement to capture how their practices, perceptions, and norms evolve over time.

 

Ecologies of Artifacts

 

AI systems rarely operate in isolation, instead functioning within interconnected platforms, interfaces, and algorithmic feedback loops (Ozmen Garibay et al., 2023). To understand how these systems interact, compete, or reinforce one another, we intentionally shifted our analytical focus away from tool-specific adoption. Instead, we examined the broader behavioral patterns—such as how freelancers disclose or conceal AI use—and universal impacts that cut across tool boundaries, such as differences in compensation or visibility on platforms. This approach allowed us to capture the relational and contextual nature of AI adoption within online labor ecosystems, where seemingly individual choices around AI use often reverberate through shared infrastructures like reputation systems and algorithmic rankings.

 

Discussion and Closing Thoughts

 

This essay argues that scale is a critical dimension for assessing foundation systems before large-scale deployment. Developing strategies to examine scale without needing large participant pools is a critical challenge for AI researchers and practitioners.

 

To this end, we build on and extend the concept of scale hacking from HCI literature, translating it for communication and media research concerned with systemic, cross-platform impacts. While prior HCI work has used scale hacking primarily to reflect on methodological adaptation, we develop it further as a bridge between design-oriented and communicative perspectives on sociotechnical systems. Specifically, we identify four dynamics that proved particularly generative in our own study: (1) triangulating methods to surface multi-stakeholder insights; (2) treating data as design material rather than static evidence; (3) envisioning technological influence over time to account for temporal scaling; and (4) analyzing ecologies of artifacts to situate AI systems within broader infrastructures. Together, these strategies allowed us to trace systemic and temporal patterns within online labor markets, offering a model for how communication researchers might study the distributed effects of large AI systems. In doing so, we position scale hacking as a cross-disciplinary framework for understanding how scale itself mediates the social life of AI.

 

This approach complements prevailing discourses in responsible AI and governance that emphasize pre-deployment evaluation. While audits, benchmarks, and red-teaming protocols can identify discrete risk categories, they are often ill-suited for capturing relational, interpretive, and socially situated harms; that is, harms that emerge through social interaction, interpretation, and power relations rather than through direct technical failure. Such harms can arise before mass adoption yet become visible only through attention to contextual dynamics. We show how mixed-method, anticipatory approaches foreground community expertise and socioeconomic positioning alongside technical assessments.

 

Beyond methodology, scale hacking challenges how HCI and AI research define scale. Instead of treating scale as larger samples or wider deployment, it emphasizes interpretability, reflexive research practice, and whose knowledge counts. This perspective aligns with calls for situated accountability (Henriksen et al., 2021; Procter et al., 2023) and encourages scholars to ask not only what changes as systems scale but who benefits—and who is left out.

 

The implications extend across stakeholder domains. For researchers, our findings underscore the value of embedding harm discovery into early design cycles, particularly for general-purpose technologies. For platform operators, the results reveal the need for clearer and more equitable norms regarding AI disclosure. For policy makers and funders, this work highlights the importance of governing not only high-risk, frontier models but also slow-moving, cumulative harms in everyday digital ecosystems.

 

Ultimately, scale hacking constitutes both a methodological intervention and a call for critical imagination. Addressing the harms of foundation models requires not only technical or regulatory innovation but also a reconfiguration of how, where, and for whom sociotechnical knowledge is produced.

 

 

References

 

Alfaras, M., Tsaknaki, V., Sanches, P., Windlin, C., Umair, M., Sas, C., & Höök, K. (2020). From biodata to somadata. Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, 1–14. https://doi.org/10.1145/3313831.3376684

 

Benjamin, J. J., Berger, A., Merrill, N., & Pierce, J. (2021). Machine learning uncertainty as a design material: A post-phenomenological inquiry. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, 1–14. https://doi.org/10.1145/3411764.3445481

 

 

Bødker, S., & Klokmose, C. N. (2012). Dynamics in artifact ecologies. Proceedings of the 7th Nordic Conference on Human-Computer Interaction, 448–457. https://doi.org/10.1145/2399016.2399085

 

Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N., Chen, A., Creel, K., Davis, J. Q., Demszky, D., . . . Liang, P. (2022). On the opportunities and risks of foundation models. https://crfm.stanford.edu/assets/report.pdf

 

Brown, B., Bødker, S., & Höök, K. (2017). Does HCI scale? Scale hacking and the relevance of HCI. Interactions, 24(5), 28–33. https://doi.org/10.1145/3125387

 

Coulton, P., Lindley, J., Sturdee, M., & Stead, M. (2017). Design fiction as world building. Proceedings of the 3rd Biennial Research Through Design Conference, 1–16. https://doi.org/10.6084/M9.FIGSHARE.4746964

 

Domínguez Hernández, A., Krishna, S., Perini, A. M., Katell, M., Bennett, S., Borda, A., Hashem, Y., Hadjiloizou, S., Mahomed, S., Jayadeva, S., Aitken, M., & Leslie, D. (2024). Mapping the individual, social and biospheric impacts of foundation models. Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, 776–796. https://doi.org/10.1145/3630106.3658939

 

Farias, P. G., Bendor, R., & van Eekelen, B. F. (2022). Social dreaming together: A critical exploration of participatory speculative design. Proceedings of the Participatory Design Conference, 147–154. https://doi.org/10.1145/3537797.3537826

 

Henriksen, A., Enni, S., & Bechmann, A. (2021). Situated accountability: Ethical principles, certification standards, and explanation methods in applied AI. Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, 574–585. https://doi.org/10.1145/3461702.3462564

 

Korsgaard, H., Lyle, P., Saad-Sulonen, J., Klokmose, C. N., Nouwens, M., & Bødker, S. (2022). Collectives and their artifact ecologies. Proceedings of the ACM on Human-Computer Interaction, 1–26. https://doi.org/10.1145/3555533

 

Kraut, R., Scherlis, W., Mukhopadhyay, T., Manning, J., & Kiesler, S. (1996). HomeNet: A field trial of residential Internet services. Proceedings of the SIGCHI Conference on Human Factors in Computing Systems Common Ground - CHI ’96, 284–291. https://doi.org/10.1145/238386.238531

 

Larsen, B. C. (2021). A framework for understanding AI-induced field change: How AI technologies are legitimized and institutionalized. Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, 683–694. https://doi.org/10.1145/3461702.3462591

 

Lee, H., & Whitley, E. A. (2002). Time and information technology: Temporal impacts on individuals, organizations, and society. The Information Society, 18(4), 235–240.

 

Moradi, F., Wiberg, M., & Hansson, M. (2018). Scaling interaction: From small-scale interaction to architectural scale. Interactions, 25(6), 90–92. https://doi.org/10.1145/3274574

 

Notess, M., & Blevis, E. (2004). Comparing human-centered design methods from different disciplines: Contextual design and principles. In J. Redmond, D. Durling, & A. de Bono. (Eds.), Futureground—DRS International Conference 2004, 17–21 November, Melbourne, Australia. https://dl.designresearchsociety.org/drs-conference-papers/drs2004/researchpapers/53

 

Ozmen Garibay, O., Winslow, B., Andolina, S., Antona, M., Bodenschatz, A., Coursaris, C., Falco, G., Fiore, S. M., Garibay, I., Grieman, K., Havens, J. C., Jirotka, M., Kacorri, H., Karwowski, W., Kider, J., Konstan, J., Koon, S., Lopez-Gonzalez, M., Maifeld-Carucci, I., . . . Xu, W. (2023). Six human-centered artificial intelligence grand challenges. International Journal of Human–Computer Interaction, 39(3), 391–437. https://doi.org/10.1080/10447318.2022.2153320

 

Procter, R., Tolmie, P., & Rouncefield, M. (2023). Holding AI to account: Challenges for the delivery of trustworthy AI in healthcare. ACM Transactions on Computer-Human Interaction, 30(2), 1–34. https://doi.org/10.1145/3577009

 

Rauh, M., Mellor, J., Uesato, J., Huang, P.-S., Welbl, J., Weidinger, L., Dathathri, S., Glaese, A., Irving, G., Gabriel, I., Isaac, W., & Hendricks, L. A. (2022). Characteristics of harmful text: Towards rigorous benchmarking of language models. Advances in Neural Information Processing Systems, 24720–24739. https://dl.acm.org/doi/10.5555/3600270.3602063

 

Sengers, P., Williams, K., & Khovanskaya, V. (2021). Speculation and the design of development. Proceedings of the ACM on Human-Computer Interaction, 1–27. https://doi.org/10.1145/3449195

 

Shelby, R., Rismani, S., Henne, K., Moon, A., Rostamzadeh, N., Nicholas, P., Yilla-Akbari, N., Gallegos, J., Smart, A., Garcia, E., & Virk, G. (2023). Sociotechnical harms of algorithmic systems: Scoping a taxonomy for harm reduction. Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, 723–741. https://doi.org/10.1145/3600211.3604673

 

Weidinger, L., Uesato, J., Rauh, M., Griffin, C., Huang, P.-S., Mellor, J., Glaese, A., Cheng, M., Balle, B., Kasirzadeh, A., Biles, C., Brown, S., Kenton, Z., Hawkins, W., Stepleton, T., Birhane, A., Hendricks, L. A., Rimell, L., Isaac, W., . . . Gabriel, I. (2022). Taxonomy of risks posed by language models. Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, 214–229. https://doi.org/10.1145/3531146.3533088

 

Wiberg, M., & Stolterman, E. (2021). Time and temporality in HCI research. Interacting with Computers, 33(3), 250–270. https://doi.org/10.1093/iwc/iwab025

Wong, R. Y., & Nguyen, T. (2021). Timelines: A world-building activity for values advocacy. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, 1–15. https://doi.org/10.1145/3411764.3445447

 

Xu, W., Dainoff, M. J., Ge, L., & Gao, Z. (2023). Transitioning to human interaction with AI systems: New challenges and opportunities for HCI professionals to enable human-centered AI. International Journal of Human–Computer Interaction, 39(3), 494–518. https://doi.org/10.1080/10447318.2022.2041900

 

Zhou, C., Li, Q., Li, C., Yu, J., Liu, Y., Wang, G., Zhang, K., Ji, C., Yan, Q., He, L., Peng, H., Li, J., Wu, J., Liu, Z., Xie, P., Xiong, C., Pei, J., Yu, P. S., & Sun, L. (2024). A comprehensive survey on pretrained foundation models: A history from Bert to ChatGPT. International Journal of Machine Learning and Cybernetics, 16(12), 9851–9915. https://doi.org/10.1007/s13042-024-02443-6

 

 

Copyright © 2026 (Angel Hwang, [email protected]). Licensed under the Creative Commons Attribution Non-commercial No Derivatives (by-nc-nd). Available at https://ijoc.org.

https://doi.org/10.65476/e20tjk66