Abstract
The emergence of text-to-image AI platforms (Flux, Midjourney, and Stable Diffusion) represents a profound shift in creative technology, yet a critical sociotechnical understanding of how proprietary architectures encode underlying sociopolitical biases remains fundamentally underdeveloped. This study utilizes the Collaborative AI Creativity Model, a novel framework for integrated social-technical evaluation across 306 generated images. Rigorous analysis reveals platform-specific interpretative “creative signatures”: Flux affords stability and utility through high technical precision and prompt adherence, while Midjourney affords aesthetic agency via metaphorical interpretation and creative expansion. Critically, the research documents a systemic divergence in subject matter handling: non-social scenes (e.g., Nature/Landscape) achieved high success, contrasting sharply with a structural failure in rendering complex Human Group Scenes (success rates dropping significantly to 18.8%–31.2%). This complexity-dependent performance degradation is interpreted not as technical error but as empirical evidence of systemic algorithmic defaults and inherent biases rooted in skewed training data. It confirms an algorithmic worldview that inherently prioritizes the processing of easily quantifiable material culture over nuanced social representation. These results, analysed through Affordance Theory, necessitate an urgent critical shift toward understanding generative AI through the prism of political economy, demanding adversarial design methodologies to mitigate embedded systemic bias in AI images.
Keywords
Introduction
The dramatic and accelerated proliferation of text-to-image artificial intelligence platforms: Flux, Midjourney, and Stable Diffusion among the most prominent, marks an important shift in visual communication and the political economy of creative labour (Cetinic and She, 2022; Laba, 2024b; Tilford, 2024). These systems translate textual descriptions into visual artefacts instantaneously. The speed is unprecedented. What began as experimental concepts have matured into commercially influential platforms within a remarkably compressed timeframe (Cetinic and She, 2022). As generative technologies infiltrate traditionally specialized professional and artistic domains, we require a nuanced understanding of their operational characteristics and interpretive patterns that extends beyond utilitarian metrics alone.
The prevailing discourse surrounding these generative capabilities is frequently anchored in what Campolo and Crawford (2020) term “enchanted determinism,” a technophilic rhetoric employed by industry advocates who champion the systems’ supposed capacity to expand human imaginative power. This rhetorical performance, which resonates with the Californian ideology (Barbrook and Cameron, 1996), strategically emphasizes innovation and efficiency. It positions technology as a resolution to complex societal challenges, thereby framing the future as a solvent for present issues (Forlano, 2021). Crucially, this framing affords commercial actors a convenient mechanism to absolve themselves of political responsibilities for the social, ethical, and representational implications of the systems they deploy (Dourish and Bell, 2014).
This utopian construction, however, obscures a significant detail: the generative process, celebrated as spontaneous creation, is fundamentally a complex social action that remains highly homogenized and proceduralized (Laba, 2024b). While these platforms yield distinctly unique outputs, they share a foundational technological blueprint. Large-scale latent diffusion models constitute the shared architecture, deep learning systems trained on colossal, often proprietary, datasets of image-text pairs (Liu and Chilton, 2021; Rombach et al., 2022; Stable Diffusion, 2022). This commonality establishes a crucial analytical premise: because the core generative mechanism is shared, the observed divergences in output cannot be attributed solely to technical novelty. Rather, they must be understood as manifestations of deliberate proprietary fine-tuning, biased training data composition, and design philosophy elements influenced by market priorities and ideological choices (Campolo and Crawford, 2020). Viewed through this lens, these foundation models transcend the status of mere technical instruments. They operate as powerful platform models with demonstrable ideological influence over creative output (Burkhardt and Rieder, 2024).
The critical research gap demanding this investigation lies not in the absence of comparative performance statistics, a utilitarian focus that risks overlooking deeper sociopolitical implications (Ananny and Crawford, 2016; Pavlichenko and Ustalov, 2023), but in the urgent requirement for an integrated framework capable of structured external observability. The phenomenon of algorithmic opacity intensifies this necessity. The internal logic of deep neural networks operates in ways largely hidden from human comprehension (Castelvecchi, 2016; Pasquale, 2015; Von Eschenbach, 2021). This structural obscurity means the system’s capacity to represent the world is constrained by perceptual biases arising from inductive biases in the machine vision system. Consequently, we must move beyond asking how platforms perform on various creative tasks. We must ask why those performance differentials matter in terms of ethical stakes and the implicit algorithmic worldviews encoded within the foundational code (Davis, 2020).
This study addresses this deficiency by deploying the Collaborative AI Creativity Model, a theoretical framework engineered specifically for integrated social-technical evaluation. The purpose of this analysis is twofold: first, to systematically quantify and compare the distinctive performance characteristics, the measurable tension between technical fidelity (prompt adherence) versus creative expansion (algorithmic agency) across Flux, Midjourney, and Stable Diffusion; and second, to utilize these empirical differences as instruments to interrogate the existence and manifestation of systemic algorithmic biases (Combs et al., 2024).
The central thesis posits that observed platform divergence is driven by distinct affordances (Nagy and Neff, 2015), which, when analysed through the CACM framework, reveal fundamental structural biases. This claim is substantiated by bifurcated empirical evidence: non-social subjects, exemplified by Nature and Landscape prompts, achieve consistently high success rates (reaching up to 61.3% for Flux). This contrasts sharply with pervasive structural failure documented in subject categories requiring complex social or anatomical representation. Human Group Scenes, for instance, show success rates dropping significantly into the 18.8% to 31.2% range across platforms (Mellamphy, 2021). We argue that this specific and predictable complexity-dependent performance degradation constitutes empirical evidence of systemic algorithmic defaults rooted in skewed training data, thereby exposing the latent sociopolitical assumptions encoded within the technical architecture of these generative systems (Luccioni et al., 2023). The disparity suggests an implicit algorithmic worldview. It rigorously prioritizes the deterministic processing of easily quantifiable natural and material elements over the indeterminate complexity of human existence and interaction.
To structure this inquiry, we must conceptualize AI image generation as a sociotechnical practice situated at the nexus of three inseparable components (boyd and Crawford, 2012; Laba, 2024b).
First, mythology: the widespread belief, often commercially reinforced, that AI systems attain authenticity, truth, and epistemic objectivity because their generative processes are rooted in disembodied machine learning, thereby establishing cultural meanings in a highly regimented technical way. Second, technology: the maximizing of computational capacity through deep neural networks, a process that operates in ways largely hidden from human comprehension, necessitating an integrated analytical framework for external observability. Third, representation: the discursive process wherein a social actor utilizes the AI medium to translate a verbal prompt into a visual output for professional, political, or ideological purposes, deliberately blurring the traditional lines defining human versus machine creative agency (Laba, 2024b; Pasquinelli, 2019).
This sociotechnical conceptualization is complemented by the integration of Affordance Theory (Davis, 2020; Gibson, 1979; Nagy and Neff, 2015), providing the necessary relational actor-system property. This approach utilizes Davis’s (2020) mechanisms (hard-coded design features) and conditions (social and cultural contexts) to explain how distinct platform architectures enable or constrain the user’s creative practice (Davis, 2023). For instance, the system architecture’s mechanism optimized for geometric stability (Flux) affords utility and low-variance output. Meanwhile, the mechanism utilizing sophisticated aesthetic filtering (Midjourney) affords aesthetic agency through metaphorical interpretation, dictating the nature and intensity of intellectual labour required from the human prompter to overcome or exploit the model’s default tendencies.
The overarching objective of this study is to advance the research agenda on critical machine vision by responding to the following integrated research inquiries: How do the performance characteristics of technical fidelity and creative expansion differ across leading generative diffusion models (Flux, Midjourney, and Stable Diffusion), and how do these measurable differences inform the creative signature of each platform? In what ways does prompt complexity affect platform performance, and does this resultant complexity-dependent degradation provide empirical evidence of systemic bias in the representation of social versus non-social subjects? How do the affordances realized through the platform’s architectural mechanisms constrain or enable the user’s creative agency? What are the resulting ethical stakes of the implied algorithmic worldview regarding human representation?
Literature review
The political economy of foundation models and platform power
Contemporary generative AI systems must be scrutinized not merely as advanced software tools but as powerful platform models with significant political and economic weight (Burkhardt and Rieder, 2024). This perspective is essential. The deployment and functionality of systems like Flux, Midjourney, and Stable Diffusion are inextricably linked to market dominance and the exertion of ideological control over creative output (Goriunova, 2013). The choice among these platforms, ranging from Flux, a closed, proprietary system prioritizing industrial speed, to Stable Diffusion, an open-source model prioritizing community flexibility, is inherently a choice among differing operational philosophies, each carrying distinct economic and political consequences (Burkhardt and Rieder, 2024; Laba, 2024b).
This critical context is magnified by the systemic perpetuation of inequality through the diffusion of machine learning systems (Campolo and Crawford, 2020). The architectural design and training of these foundation models inevitably embed and amplify societal biases, a phenomenon extensively documented in critical data studies (Buolamwini and Gebru, 2018; Bender et al., 2021; Crawford, 2021; Kim et al., 2023). Studies on Large Language Models that inform the prompt-processing functionality in visual generative systems reveal alarming evidence of perpetuating regressive gender and racial stereotypes (Bianchi et al., 2023). This bias is not a theoretical risk. It is a measurable empirical reality. The outputs of AI image generators frequently construct a “world according to Stable Diffusion… run by White male CEOs,” while associating minority groups with stereotypical or lower-wage roles (Nicoletti and Bass, 2023). Viewed through this lens, the outputs of these technical systems are fundamentally linked to the unequal distribution of cultural and political representation present in their colossal training data. Consequently, the necessity of an external, structured framework, such as the Collaborative AI Creativity Model, is validated as a means to monitor and quantify output bias, regardless of whether the underlying model is proprietary or open source.
The technical challenge: From utility metrics to algorithmic opacity
A substantial body of emergent research originating primarily from computer science and human-computer interaction has centred on advancing practitioner knowledge of text-to-image generation, focusing predominantly on optimizing output (Liu and Chilton, 2021; Pavlichenko and Ustalov, 2023). This utilitarian focus seeks technical precision; maximizing the system’s capacity to render accurate representations, enhance aesthetic quality, and achieve cohesive visual results, often through experimental taxonomy of prompt modifiers and model hyperparameters (Oppenlaender, 2023). While instrumental in mapping the operational limits and possibilities of human-model interaction, this approach risks overlooking the deeper sociopolitical implications inherent in design choices (Ananny and Crawford, 2016).
The transition from utility-centric assessment to critical sociotechnical critique is complicated by the fundamental issue of algorithmic opacity (Pasquale, 2015). Foundation models, relying on deep neural networks, operate in covert and often incomprehensible ways. Growing concerns suggest that opaque systems harbour biases that remain undetected (Castelvecchi, 2016; Von Eschenbach, 2021). No analytical framework can provide genuine internal transparency into proprietary architecture (Laba, 2024b). Therefore, the critical gap lies not in achieving internal transparency which is often impossible in commercial “black box” systems, but in developing a methodology capable of facilitating structured, external observability of the internal creative process. This transforms subjective output evaluation into empirical data suitable for social critique (Laba, 2024b). This structured approach allows researchers to isolate and categorize the patterns of algorithmic decision-making, which traditional holistic models are incapable of penetrating (Ananny and Crawford, 2016; Laba, 2024b).
Machine originality and the rhetoric of aesthetic mimicry
The debate surrounding the interplay of machine originality and repetition represents a central theoretical tension in the critical appraisal of generative AI. Traditional concepts of creativity, rooted in romantic ideals of human experience and ascribed mind (Messingschlager and Appel, 2023), contrast sharply with the outputs of deep neural networks. These are products of deterministic, sequential computation (Laba, 2024b). Critical perspectives argue that the celebrated aesthetic brilliance of machine originality emerges fundamentally from repetition (Goriunova, 2013).
When generative models produce images, the mechanism relies not on direct copying of training data but on searching for and replicating statistical patterns in vast image-text datasets (Pasquinelli, 2019). This procedural imitation results in what is conceptually termed aesthetic mimicry (O'Meara and Murphy, 2023), where the system operates by emulating aesthetic qualities of existing works on which it has been trained. Accordingly, Tilford (2024) characterizes the output as a “synthetic imagination” that mimics established cultural products, granting users a form of artistic subjectivity previously unattainable due to its practical complexity. For Zeilinger (2021), this mechanism positions T2I generators as “generative adversarial copy machines” capable of both conformity to and subversion of established creative norms. They challenge the cultural logic of intellectual property. The practice of human-model interaction raises complex questions regarding the distribution of agency and the extent to which generative models serve as facilitators or arbiters of originality, thereby highlighting the unresolved tension between innovation and imitation inherent in this sociotechnical practice (Laba, 2024b; Messingschlager and Appel, 2023). It is precisely this ambiguity, where output is often seen as both novel and imitative, that mandates a framework capable of quantifying and locating the machine’s creative action relative to human instruction.
Theoretical framework: Affordance actualization and the CACM
To move beyond the limitations of holistic output assessment and the opacities of internal machine processes, this study integrates two critical frameworks: Affordance Actualization Theory, which grounds the human-machine relationship, and the Collaborative AI Creativity Model, which operationalizes the empirical measurement of that relationship. This dual theoretical deployment is necessary to transform the philosophical question of “what is AI creativity?” into the measurable, empirical question of “where and how does the machine’s creative action occur relative to human instruction?” (Laba, 2024b).
Affordance Actualization Theory: Mechanisms, conditions, and practice
The overarching theoretical framework for this sociotechnical evaluation is Affordance Actualization Theory, which provides a robust conceptual lens for addressing the dynamic, relational property between an actor (the human prompter) and a technological system (the generative platform) (Bao et al., 2023; Strong et al., 2014). Drawing on Gibson’s (1979) foundational work, AAT is concerned with how users perceive and utilize affordances, the possibilities for action signalled by a technological environment, to produce desired outcomes (Nagy and Neff, 2015; Strong et al., 2014). The process involves two crucial stages: Affordance Perception, which is shaped by the user’s goals and technical capabilities, and Affordance Enactment, where the perceived opportunities are acted upon to achieve an outcome (Bernhard et al., 2013).
The application of this framework for refined sociotechnical critique necessitates the integration of Davis’s (2020) mechanisms and conditions framework (Davis, 2023). First, mechanisms: these represent the hard-coded, intrinsic design features of the platform, encompassing architectural choices and internal algorithmic processes. Examples include Midjourney’s proprietary bias toward metaphorical interpretation or Flux’s optimization for geometric stability. These mechanisms actively define the system’s native interpretative tendency. Second, conditions: these comprise the necessary social, cultural, and institutional contexts required for the mechanism to function effectively. Conditions include the user’s technical skill, prompt engineering knowledge, and the established legal or ethical legitimacy of the desired output.
The relational actor-affordance property is central: a human prompter is conceptualized as a goal-oriented actor who perceives a style modifier or prompt parameter as an affordance to be acted upon to achieve certain stylistic outcomes (Laba, 2024a, 2024b). However, a core finding in the literature is that affordance actualization in visual generative media does not guarantee the desired visual outcome, owing to the inherent technological complexity and model uncertainty (Ananny and Crawford, 2016; Combs et al., 2024). This failure to actualize is critical. It requires the human actor to engage in significant intellectual labour (e.g., complex negative prompting or extensive experimentation) to overcome the platform’s default interpretative tendency, confirming that the platform’s affordances actively shape and constrain the human-AI collaborative process (Laba, 2024b).
Operationalizing machine creativity: The tension between sequential logic and creative messiness
The conceptualization of machine creativity within the CACM framework warrants detailed methodological justification, specifically addressing the concern that treating creation as a measurable sequence might overlook the inherent “messiness” and non-linearity observed in human artistic practice.
The highly structured, five-stage protocol of the CACM is deployed as a deliberate methodological decision, rather than an ontological claim about the nature of creativity itself. While human creativity is recognized as non-linear, iterative, and inherently subjective, machine creativity, based on current LDM architectures, is a functional product of deterministic, sequential computation. The CACM exploits this mechanistic nature. It systematically breaks down the process from Human Creative Input to Output Evaluation, allowing the researcher to isolate the specific point where the machine’s internal logic deviates from human intent. By providing structured, external observability of this computational process, the CACM framework furnishes the necessary analytic tool for social scientists to evaluate black box systems methodically and systematically. This approach transforms the analysis of subjective attributes, such as “creative interpretation” and “emotional tone,” into quantifiable, empirical data suitable for robust social critique, as demonstrated by the achievement of a high intercoder reliability score (kappa = 0.82) in the study’s data protocol.
The Collaborative AI Creativity Model (CACM): Defining the creativity-error distinction
The Collaborative AI Creativity Model serves as the theoretical and operational foundation for this comparative analysis, transitioning the study of generative AI from holistic outcome assessment to structured, five-stage process evaluation of human-machine interaction. The framework provides several theoretical advancements over previous models, which often either treated AI creativity as a purely automated technical process or as an unmeasurable conceptual “black box” operation (Elgammal et al., 2017; Zeilinger, 2021).
The CACM’s central methodological innovation is its capacity to enable the crucial distinction between technical execution failure (error) and intentional algorithmic deviation (creative agency). This distinction is achieved through the definition of five interconnected components, allowing researchers to gain quantifiable data points necessary to locate and measure algorithmic influence across different stages of image generation. First, Human Creative Input: this initial stage systematically examines how creative intent is encoded through prompts, focusing on parameters such as complexity, subject matter, style, and emotional tone. Second, AI Interpretation (Fidelity/Error): this component assesses the machine’s literal execution of the prompt, measured by Prompt Adherence (Low, Moderate, or High). A Low score here (e.g., anatomical distortion, missing key elements) is categorized as a technical execution failure or error, revealing the system’s structural limits. Third, Creative Expansion (Agency/Deviation): this stage measures the machine’s capacity to elaborate beyond the prompt’s explicit text, quantified by Novel Element Introduction (None, Minor, Significant) and Unexpected Interpretations. A High score is categorized as a creative departure or agency, provided the output maintains overall relevance to the initial intent. To distinguish between “Error” and “Creative Expansion,” we applied a criterion of Semantic Coherence. A “Novel Element” was coded as an act of Agency only if it maintained physical logic and narrative consistency (e.g., adding ambient lighting or contextually relevant objects like a vase on a dinner table). Also, elements that breached anatomical or physical logic (e.g., a hand merging into a table or floating, disconnected limbs) were strictly coded as “Technical Errors,” regardless of the prompt’s complexity. This distinction ensures that the study does not conflate the model’s “hallucinations” or failures in spatial binding with intentional creative variation. Fourth, Aesthetic Realization: this component incorporates technical metrics of output quality, including overall visual quality, colour harmony, and stylistic consistency. Fifth, Output Evaluation: the final component provides a holistic assessment synthesizing prompt alignment, the creativity-fidelity balance, and the final image’s uniqueness.
By systematically isolating fidelity from agency, the CACM framework transforms the philosophical question of “what is AI creativity?” into a robust, measurable empirical inquiry necessary for conducting deep sociopolitical critiques of algorithmic design.
Methodology and data protocol
This study employed a mixed-methods design, strategically combining quantitative analysis of performance metrics (fidelity and agency) with qualitative assessment of creative interpretation and algorithmic bias, in alignment with the demands of sociotechnical critique (Davis, 2020). The core objective was comparative analysis of three leading generative diffusion platforms; Flux (FLUX 1.1), Midjourney (Version 6.1), and Stable Diffusion (3.5 Large), utilizing a meticulously constructed dataset of 102 systematically varied prompts. This standardized prompt set facilitated comprehensive analysis, yielding a total dataset of 306 generated images (102 prompts per platform) for subsequent coding and evaluation. The entire data collection was executed within a tightly controlled temporal window, occurring between September and November 2024. We utilized standardized default settings across all three systems to ensure that output differences could be robustly attributed to architectural variance rather than user parameter manipulation.
A crucial point of technical grounding that informs the critical thesis is the explicit clarification that all three platforms rely fundamentally on latent diffusion models as their core generative technology (Rombach et al., 2022; Stable Diffusion, 2022). This establishes the definitive premise that observed divergences in output are not rooted in fundamentally distinct generative approaches (e.g., GANs vs Diffusion), but rather in proprietary architectural fine-tuning, training data curation, and optimization for distinct market segments (Burkhardt and Rieder, 2024; Flux, 2023). The resultant performance differentials are thus interpreted as ideological and commercial choices, rather than immutable technical constraints. This justifies the sociopolitical critique advanced by this research (Campolo and Crawford, 2020).
The data generation adhered to a stringent protocol designed to mitigate inherent sampling biases prevalent in generative AI research. The 102 prompts were systematically categorized based on two axes: Complexity (Simple: 1–2 elements; Complex: 5+ elements) and Subject Matter (Social/Human-Centric vs Non-Social/Material). This structured variation was essential for isolating the variable of complexity-dependent performance degradation, thereby furnishing the empirical evidence required to test the thesis concerning systemic bias in human representation (Combs et al., 2024).
Midjourney, due to its default function, generates four images per prompt output. To eliminate potential selection bias, where a researcher might manually choose the most aesthetically pleasing or most prompt-adherent image, and to ensure that the sampled output reflects the platform’s default world model rather than a human-curated “best” result, the first image from each output was consistently selected for evaluation. This methodological consistency reinforces the study’s objective of scrutinizing the system’s intrinsic interpretative tendency as defined by its mechanisms and conditions (Davis, 2020).
To ensure the methodological precision necessary for assessing the subjective categories within the CACM framework, specifically Prompt Adherence, Aesthetic Realization, Creative Expansion, and Emotional Tone; robust measures for intercoder reliability were implemented. This process directly addresses methodological scepticism regarding the objectivity of creative evaluation in social science.
A subsample of 30 images (10 per platform) was independently evaluated by three experienced coders, all possessing expertise in visual communication and art studies. The quantitative metrics achieved a high degree of concordance, resulting in a Cohen’s Kappa score of kappa = 0.82. For the more nuanced, subjective categories, a consensus rate of 91% percentage agreement was established. The reliability process incorporated scheduled structured discussion sessions focused specifically on resolving ambiguous coding cases. These calibration sessions ensured consistent application of CACM criteria, critically refining the distinction between a Minor novel element (an acceptable enhancement) and a Significant departure (genuine algorithmic agency), thereby isolating creative deviation from technical execution failure. The demonstrated ability to consistently measure these subjective attributes validates CACM as a robust analytic tool capable of transforming the philosophical critique of algorithmic influence into quantifiable, empirical data suitable for social and humanist analysis (Ananny and Crawford, 2016; Llano et al., 2022).
Results
The quantitative and qualitative analysis of the 306 generated outputs reveals distinct platform-specific interpretative tendencies rooted in an architectural trade-off between technical fidelity and creative agency, quantified precisely through the integrated framework of the Collaborative AI Creativity Model. The empirical data is systematically organized to confirm the initial theoretical propositions regarding platform affordances, complexity-dependent degradation, and the systemic manifestation of algorithmic bias within subject matter handling.
Platform performance: Fidelity, creativity, and interpretative tendencies
Prompt Adherence (High) versus Creative Expansion (High).
Flux consistently demonstrated the highest capability for literal execution, achieving 45.1% High Adherence. This outcome suggests that its proprietary architectural mechanism is fundamentally optimized to prioritize low-variance, reliable translation of human intent, consequently affording stability and precision to the prompter. Conversely, Midjourney executed a systematic architectural decision, achieving the highest rate of High Creative Expansion (42.2%) at the demonstrable cost of a marginally lower prompt adherence rate (40.2%). The minimal 4.9% difference in prompt adherence between Flux and Midjourney corresponds to a substantial increase in outputs characterized by high creative license. This quantitatively confirms the architectural prioritization of aesthetic departure over strict fidelity. Stable Diffusion, even utilizing its advanced large version, consistently exhibited the lowest scores in both adherence (31.4%) and high creativity (22.5%), positioning it as a low-agency system whose architecture tends toward cautious, technical literalism constrained by inherent execution limitations.
Creative expansion: Novel Element Introduction and algorithmic agency
Creative expansion, the operational measure of algorithmic agency, was meticulously assessed through the introduction of non-requested, contextually relevant elements (Novel Elements) and interpretive deviation (Unexpected Interpretations).
Novel element introduction rates.
Complexity-dependent performance degradation
To assess architectural robustness, the relationship between prompt complexity (Simple: 1–2 elements; Complex: 5+ elements) and reliable output was quantitatively measured.
High adherence rate degradation by prompt complexity.
Critical subject matter analysis and algorithmic defaults
The most significant empirical evidence supporting the study’s critical thesis that architectural design reflects and amplifies sociopolitical biases emerged from the categorical analysis of subject matter handling, specifically the pervasive discrepancy between non-social and human-centric scenes. Success was defined by achieving a “High (excellent composition, very appealing)” rating in the Aesthetic Realization component.
High quality output success rate by subject category.

1. Flux (a), Midjourney (b), and Stable Diffusion (c) (in the order of appearance)-Prompt (“A peaceful mountain stream with crystal clear water flowing over rocks”).
Flux excelled in this domain, leveraging its geometric precision for accurate landscape modelling. Qualitative analysis reveals that Flux’s nature scenes demonstrated superior environmental coherence, with a nuanced handling of light-atmosphere interactions. Midjourney’s nature outputs, while technically proficient, showed a tendency toward more dramatic and stylized interpretations, often enhancing natural elements beyond strict realism.
Urban and architectural scenes revealed more pronounced differences between platforms. Flux achieved 52.4% high-quality ratings, with notable strength in geometric accuracy and perspective consistency. Midjourney followed at 47.6%, excelling in atmospheric urban scenes but occasionally struggling with precise architectural details. Stable Diffusion showed more limited success at 33.3%, with challenges in complex urban environments. The qualitative assessment (Figure 2(a)–(c)) indicates that while Flux prioritized structural accuracy, Midjourney often introduced creative elements that enhanced the scene’s mood but sometimes compromised architectural precision. 2. Flux (a), Midjourney (b), and Stable Diffusion (c) (in the order of appearance)-Prompt (“A tall building seen from below, with the sky distorted by its height”).
Second, social failure (the crisis of representation): human representation consistently proved the weakest category across all platforms. Crucially, performance plummeted when complex social interactions were required. Human Group Scenes showed the lowest success rates (ranging from 3. Flux (a), Midjourney (b), and Stable Diffusion (c) (in the order of appearance)-Prompt (“A large family gathered around a dinner table, laughing and sharing stories”). 4. Flux (a), Midjourney (b), and Stable Diffusion (c) (in the order of appearance)-Prompt (“A joyful crowd celebrating at a summer festival”).

Midjourney demonstrated superior Overall Visual Quality (High:
Flux exhibited notable Technical Execution consistency (Excellent:
Discussion
Interpreting platform affordances: The trade-off between fidelity and agency
The quantified architectural trade-off between fidelity (prompt adherence) and agency (creative expansion) is precisely explained by grounding the performance metrics in Davis’s (2020) mechanisms and conditions framework of Affordance Actualization Theory (Davis, 2023). The unique mechanisms of each platform actively define the spectrum of affordances available to the user, thereby constraining or enabling specific creative practices.
Flux: Affording utility and precision
Flux’s consistently high prompt adherence (45.1% High Adherence) and its minimal degradation rate (21.0% drop from simple to complex prompts) are the direct manifestation of a specific architectural mechanism: proprietary fine-tuning optimized for geometric stability and low-variance, literal interpretation (Davis, 2023). This design choice, consequently, affords utility and precision for the human actor (Nagy and Neff, 2015). The structurally coherent outputs for Urban/Architectural scenes (Figure 2(a)) visually confirm this affordance. Flux intentionally constrains the degree of creative expansion, minimizing the risk of “unexpected interpretations” to ensure reliability essential for professional workflows where precise, low-variance image generation is required.
Midjourney: Affording aesthetic agency and metaphorical interpretation
Midjourney, characterized by its lead in High Creative Expansion (42.2%) and the rate of Significant Novel Elements (35.3%), operates via a contrasting mechanism: sophisticated aesthetic filtering coupled with deep metaphorical weighting in its prompt processing. This mechanism explicitly affords aesthetic agency and creative deviation (Laba, 2024b). For artists, illustrators, and conceptual designers, Midjourney offers a powerful tool for generating visually striking results that enrich the prompt with atmospheric mood and interpretive layering (O'Meara and Murphy, 2023). However, this agency comes with a crucial condition: the user must accept a higher degree of algorithmic deviation where the output may prioritize the system’s inherent aesthetic worldview over strict literal prompt adherence (Pasquinelli, 2019). The data explicitly validates this trade-off. Midjourney achieves superior overall visual quality (52.9% High Quality) by strategically leveraging creative expansion, confirming that its technical excellence is fundamentally linked to its capacity for artistic departure (Kittler, 2006).
Stable Diffusion: Affording openness and configurability
Stable Diffusion’s performance profile, the lowest adherence (31.4%), lowest novelty introduction (18.6% Significant Novel Elements), and the highest complexity degradation (43.8% drop), must be interpreted through its unique political economy affordance: openness and configurability (Burkhardt and Rieder, 2024; Stable Diffusion, 2022). Its core mechanism is an open-source architecture that, by its very nature, lacks the tightly controlled, proprietary aesthetic fine-tuning of its competitors. This structural decision affords unparalleled user flexibility and customization; the ability to fine-tune the model to overcome inherent defaults, but demands significant technical literacy from the user (the condition) to achieve high-quality, complex output (Davis, 2023). When operating at default settings, Stable Diffusion often resorts to technical literalism (Table 2), as its mechanism is optimized for generalized reconstruction rather than subjective aesthetic polish. This signature reflects a model designed as a foundational technology for research and development rather than a commercially polished end-user product like Midjourney.
The algorithmic worldview: Complexity, bias, and the crisis of human representation
The stark, systemic failure documented in the Human Group Scenes category (success rates dropping to 18.8% for Stable Diffusion and 25.0% for Flux) provides the definitive empirical foundation for interrogating the ethical and political stakes of these platforms (Burkhardt and Rieder, 2024). This bifurcation reveals the implicit algorithmic worldview encoded within the architecture.
The mechanism of complexity-dependent bias facilitation
The comparison between the robust proficiency in generating non-social subjects (Nature/Landscape success rates up to 61.3%) and the pervasive failure in complex human scenes reveals a dual causality: architectural limitation and data bias. Technically, the struggle to render group scenes is often attributed to the “binding problem” in Latent Diffusion Models, where the system’s cross-attention mechanisms fail to correctly associate specific semantic attributes (such as hands or facial features) with distinct entities in a crowded spatial composition (Rombach et al., 2022). However, we argue that these architectural constraints do not operate in a vacuum; they actively expose the sociopolitical skew of the training data. When the model encounters high computational complexity and fails to resolve spatial coherence, it defaults to the most statistically probable patterns found in its dataset, patterns that are demonstrably skewed toward material culture and stereotypical social representations (Bianchi et al., 2023).
Consequently, complexity-dependent performance degradation (Table 3) functions as a mechanism of bias facilitation where technical uncertainty forces a reliance on dataset defaults (Davis, 2023). When the prompt demands high relational complexity (e.g., “A large family gathered around a dinner table”), the system strains the limits of its data model. As vividly illustrated by the failures in Figures 3 and 4, the result is often anatomical failure, spatial incoherence, or the application of homogeneous, statistically overrepresented figures when diversity or complexity is implied. This mechanism directly addresses the critical question of why generated images often consistently render white, anatomically simplified, or stereotypical figures when diversity or complexity is requested (Laba, 2024b; Nicoletti and Bass, 2023). The technical failure (e.g., anatomical distortion, missing figures) is the manifestation of social bias embedded in the input data (Mellamphy, 2021). The system defaults to the path of least technical resistance. This tragically often correlates with demographic overrepresentation and the omission of social complexity.
Ethical stakes in platform architectures
This critical analysis underscores the political economy inherent in these platform models (Burkhardt and Rieder, 2024). Flux’s high prioritization of precision, while technically successful in controlled scenes, must be critically examined for how its low-variance design choices risk reinforcing existing social defaults. If a platform is optimized for predictable adherence, it may be less likely to generate novel or diverse human interpretations if the source training data is historically and socially biased. Conversely, while Midjourney offers superior aesthetic agency, its high failure rate in complex social scenes demonstrates that its artistic mechanisms are insufficient to overcome the fundamental, structural biases of the underlying LDM regarding human representation. The collective struggle across all platforms in depicting complex human dynamics confirms the post humanist tendency of these systems to prioritize easily quantifiable material elements over the nuanced representation of diverse human existence (Mellamphy, 2021).
Rethinking machine creativity: The CACM as a tool for social critique
The implementation of the CACM framework successfully addressed the initial theoretical tension regarding its mechanistic structure. By systematically isolating technical adherence (fidelity) from creative expansion (agency) (Table 1; Table 2), the framework proved essential for distinguishing true algorithmic agency from mere technical error, a critical distinction required for a sociopolitical critique.
Locating algorithmic deviation
The sequential structure of CACM, the deliberate operationalization of creativity as a step-by-step process, allows researchers to locate the precise point of algorithmic deviation where the machine’s internal logic overtakes human intent (Llano et al., 2022). This ability to differentiate error (e.g., anatomical failure in Figure 3(c)) from agency (e.g., metaphorical enhancement in Figure 1(b)) transforms subjective evaluation into quantifiable data. It makes the “black box” operation externally observable for structured social critique, as defined in our refined Transparency Claim.
Creative signatures as operational philosophies
The term “creative signatures” is empirically validated by the CACM findings as a descriptor for platform-specific interpretative tendencies (Natale and Henrickson, 2022). These signatures are not arbitrary aesthetic choices but are reflections of distinct operational philosophies that dictate the nature of the human-AI collaboration (Atkinson and Barker, 2023). Flux’s Signature: utility, precision, and geometric coherence, reflecting an engineering-centric philosophy optimized to mitigate spatial incoherence. Midjourney’s Signature: metaphorical depth, high aesthetic polish, and creative deviation, reflecting an artistic-centric philosophy that effectively masks technical limitations through heavy aesthetic stylization. Stable Diffusion’s Signature: open-source flexibility, technically constrained by the “binding problem” inherent to raw Latent Diffusion architectures, reflecting a research- and community-driven philosophy that prioritizes access over proprietary fine-tuning.
This reinforces the argument that optimal creative outcomes are achieved not by seeking a single “best” AI but through strategic platform selection guided by a critical understanding of these unique, measurable interpretative tendencies.
Implications and future directions: The call for adversarial design
The empirical evidence necessitates a fundamental shift in how the capabilities of foundation models are assessed. Relying solely on overall quality metrics or abstract concepts of creativity fails to capture the ethical implications of complexity-dependent degradation. Our findings demonstrate that current systems are optimized for material culture processing, and their limitations in social representation are structural, rooted in the political economy of their training data collection.
Future research must prioritize two critical pathways. First, longitudinal studies tracking the evolution of these platform affordances are required to determine if proprietary architectural fine-tuning can successfully mitigate the systemic biases inherited from the underlying foundation models. Second, this study validates the urgent need for adversarial design methodologies aimed at actively challenging algorithmic defaults. Future work should systematically test prompts that explicitly demand diverse, complex social compositions to track where and how platforms fail, thus driving accountability toward ethical and socially representative outputs. The CACM framework provides the structured methodology necessary for such ongoing, social-technical auditing, ensuring that the habitual use of these technologies does not “warp or eclipse our values and our goals” (Boddington, 2021).
Conclusion
This comprehensive, sociotechnical analysis confirms that machine creativity in visual generative media is demonstrably non-uniform, driven by platform-specific interpretative tendencies rooted in distinct, measurable architectural affordances. The successful deployment of the Collaborative AI Creativity Model rigorously integrated technical performance metrics with high-level sociotechnical critique, validating its utility as a novel framework for the critical auditing of algorithmic bias and platform power in complex foundation models.
The empirical findings reveal a profound and measurable architectural trade-off that defines the operational philosophies of the leading platforms. Flux’s architecture prioritizes utility and precision, achieving the highest rates of prompt adherence and minimal complexity-dependent degradation. This affordance offers stability and predictability, making it optimal for tasks requiring high technical fidelity and low aesthetic deviation. Midjourney’s architecture prioritizes aesthetic agency and metaphorical interpretation, achieving the highest rate of creative expansion and superior overall visual quality. This affords artistic latitude but requires the user to accept a higher degree of inherent algorithmic deviation.
The analysis compellingly reinforces the argument that optimal creative outcomes are achieved not by seeking a single “best” AI but through strategic platform selection guided by a critical understanding of which measurable affordance, utility versus artistic exploration, best aligns with the project’s specific requirements. Most significantly, the empirical data revealed the profound ethical and sociopolitical implications of platform design. The consistent, complexity-dependent failure in rendering complex human representation (Human Group Scenes: 18.8%–31.2% success across platforms) reveals the systemic limitations and biases inherent in the training data’s algorithmic worldview. This measurable structural failure in representing social culture, despite proficiency in material culture, functions as the mechanism of bias facilitation. This finding dictates the urgent necessity for developers to address the ethical defaults and systemic biases encoded in training data, particularly concerning demographic diversity and social complexity, as this failure directly facilitates the reproduction of existing representational inequalities.
The originality of AI-generated content is therefore framed as a cline of human control, where the final output’s originality rests ultimately on the system’s ability to detect statistical patterns in the datasets on which it has been trained. The disproportional power rests with the platform and its creators, who determine the boundaries of ethical and aesthetic possibility.
Footnotes
Acknowledgements
We acknowledge the use of Grammarly software (pro) for fine-tuning and clarity enhancement of the English academic prose in this final manuscript. All 306 images used in this study were generated by the authors using their respective platform accounts during the specified data collection window.
Authors’ contributions
Author 1: Conceptualization, methodology, investigation, and writing – original draft.
Author 2: Data curation, formal analysis, and guidance. All authors read and approved the final manuscript.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data Availability Statement
Limitations
The study’s evaluation was conducted within a specific temporal window, capturing platform capabilities at a particular point in their rapid evolution, which may not reflect current or future capabilities. The analysis of 102 prompts per platform, while substantial, may not capture all use cases or creative scenarios. The acknowledged subjective nature of creative evaluation, even with the structured CACM framework and high intercoder reliability (kappa = 0.82), introduces potential bias in assessment. Additionally, the focus on three specific platforms limits the generalizability of findings to other systems.
