Abstract
This article investigates the emergence of epistemic uncertainty in computer vision (CV), an AI subfield that equips machines with visual capabilities. Drawing on 9 months of ethnographic fieldwork in an AI laboratory, I trace the knowledge production in CV models across training, validation, and review processes, identifying three key sources of uncertainty: flashing data affordances, ambiguous validation standards, and contested knowledge translations. To address these “invisibility problems,” scientists “operate” on image data, transforming raw image datasets into entities through sensory rather than purely quantitative metrics. By conceptualizing machine learning as a sensory, interactive, and processual knowledge system, this paper highlights the role of visual communication in shaping the epistemological foundations of AI.
Keywords
The history of artificial intelligence is, fundamentally, a history of efforts to equip machines with human-like sensory capabilities (Dobson, 2023). Computer vision (CV), known as an AI subfield, involves tasks such as acquiring, processing, generating, and understanding digital images. It enables machines to extract high-dimensional data from the real world, converting it into digital or symbolic information akin to human visual perception (Horn, 1986; Klette, 2014; Morris, 2003). While CV has excelled in applications such as facial recognition, autonomous driving, and artistic creation, it has yet to gain universal recognition as a source of theoretical knowledge with certainty (Ananny and Crawford, 2018), transparency, and robustness (Boukhelifa et al., 2017; Boukhelifa and Duke, 2009; Soden et al., 2020). For example, in generative tasks, identical prompts can produce entirely different scientific images, heightening concerns about the reliability of AI-generated knowledge and its alignment with human intentions (Christian, 2020; Ji et al., 2024). This raises an ongoing epistemological debate: can machine learning fully mechanize human vision, transforming it into a verifiable algorithmic model?
At the heart of this puzzle lies the nature of human vision. For humans, “seeing” involves both sensory experience and cognitive reflection—a dual process where information is simultaneously acquired and interpreted (Ihde, 2001; Merleau-Ponty, 1969). When we see a red traffic light, we not only perceive the color red (sensory experience) but also instantly understand its meaning as “stop” (cognitive reflection). However, in machine learning, only post hoc reasoning can be encoded, while the sensory aspect of “seeing” is often overlooked. The traditions of the history of science (e.g., Daston and Galison, 2010; Rossi, 2019) and laboratory ethnography (e.g., Alač, 2023; Cetina, 1999; Latour et al., 1986; Star, 1983) reveal that the suppression of subjective sensory experience merely reflects a positivist strategy, failing to capture the full complexity of laboratory life. Crucially, the knowledge production process, based on image data and GPUs, 1 cannot erase the “techniques of the observer” (Crary, 1992). Thus, instead of viewing machine learning as the compilation of model training results, I conceptualize CV as a form of processual knowledge—the operational image between human sensory perception and computational infrastructure.
This study examines how AI scientists train CV models with their eyes. Since December 2023, I have conducted 9 months of ethnographical fieldwork in a CV lab in Singapore, observing how a CV model was trained from the ground up. During the first 6 months, I closely followed the researchers’ experimental workflows, group discussions, and paper writing processes. To complement these observations, I analyzed project papers, source code, and presentation materials in detail. Monthly follow-ups from May to October allowed me to document their post hoc reflections and rebuttal strategies during peer review sessions, until the research was accepted by a top international AI conference. Returning to the site of knowledge production, I aimed to capture and preserve the fleeting visual traces of their work. My focus was on the translation between images, video, text, and code within the integrated development environment (IDE), illustrating what remains (in)visible to researchers throughout the model training process.
I argue that human vision cannot be fully digitized into algorithms; instead, it underpins the training of AI models. In the lab, image data are no longer directly accessible; they appear as “ghosts” on the screen. When images become invisible, researchers often express frustration with the difficulties of interpreting experiments and validating knowledge. This invisibility fragments the experimental process into granular workflows, disrupts the “hypothesis–observation–verification” cycle, and frequently demands iterative redesigns of model structures. Yet, researchers do not passively endure these challenges. They actively employ self-enhanced and socially mediated visual strategies, transforming raw image datasets into operational entities through sensory experience rather than purely quantitative metrics. As a result, making CV models resembles crafting works of art more than solving mathematical problems.
Literature review
Defining images with data
Images today are defined less by what they depict than by their entanglement with data (Anderson, 2017). In an era where machine vision and algorithmic processing saturate everyday infrastructures, we may confuse that a photo on the phone, a heatmap in a scientific paper, and a tensor in a neural network may share no visual resemblance, yet all three are referred to as “images.” One of the most epistemologically consequential shifts in contemporary media is the destabilization of the image's conceptual boundaries, triggered by the proliferation of visual forms that are inherently computable (MacKenzie and Munster, 2019). Images can be generated, parsed, and circulated by algorithms, and embedded in databases for knowledge production and verification (Coopmans et al., 2014). To address this complexity, a theoretical intervention is necessary to reclarify “what counts as an image” case by case. Thus, I propose a brief typology based on the image–data relationship (Table 1), which serves to precisely locate the research object of this study and to disentangle the divergent ontological and epistemological commitments that the concept of the “image” has acquired.
A typology of images based on image–data relationship.
Image as modality, at the most fundamental level, reconfigures the visual object into a distinct computational modality where its digital materiality takes precedence over its visual content (Morris, 2003). Here, the image is ontologically grounded as a structured array of pixel values, metadata, and tensors for computation. This perspective assumes that the image is a discrete and high-dimensional dataset and functions as a raw resource that holds weight and texture within a file system (Drucker, 2014; Manovich, 2001). For researchers in digital humanities or computer science, this material substrate is the primary site of inquiry. As theoretical work by Berry (2016) and Pink et al. (2016) suggests, treating the image as a numerical representation allows it to be subjected to the same algorithmic logic as text or audio, which fundamentally alters its status from a cultural artifact to a calculable unit of information.
Moving from raw material to representative function, image as visualization positions the image as an epistemological bridge between visual cognition and abstract data. In this context, the image is secondary to the data it serves to illuminate because its purpose is to translate the invisible complexity of statistics into the visible clarity of patterns (Gross & Harmon, 2013: 20–47). The underlying commitment here is that validity resides in the numbers while the image functions as a rhetorical device to make that truth accessible. Whether in scientific modeling or GIS mapping, the image is judged by its transparency and its covariation to the source (Tufte, 2001). The image, here, serves as a subservient tool that transforms quantitative density into sensory insight and provides the necessary evidence for human reasoning to grasp vast scales of information (Halpern, 2015; Kitchin, 2014).
Image as interface pushes the relationship further, where the visual object acquires agency and becomes an operational force. Beyond simple representation, these images act as a “programmed object” that mediates the feedback loop between human intent and the machine's database (Gaboury, 2021). This perspective assumes that the image is an actant capable of executing code so that to interact with the image is to intervene in the data structure itself (Chun, 2011). As seen in the control systems of planetary rovers or the GUI of a smartphone, the image constructs a specific environment of affordances. In the cases of space images from Mars (Vertesi, 2014) and various interfaces (Galloway, 2012), these images do not merely depict a reality but actively shape the parameters of what can be known and done by effectively collapsing the distance between the user’s eyes and the algorithm logic.
Interestingly, in the context of scientific images, the logic of interface covers mechanical computation and static representation. Historically, images function as machine-driven 2 and “fixed” visual evidence, underpinning certainty, objectivity, and consensus in scientific knowledge (Amann and Cetina, 1988; Daston and Galison, 2010). However, the subject–object dichotomy within this scientific vision remains largely an unfulfilled belief. Scholars in Science and Technology Studies (STS) and Human–Computer Interaction (HCI) have revealed that scientific image-making is far from being a precise copy or replication of objects (Coopmans et al., 2014). Instead, it is laden with the “interpretive flexibility” of researchers (Doherty et al., 2006), interacting frequently with common sense and expertise across various disciplines to create what Goodwin (1994) terms “professional vision.” Latour et al. (1986) introduces the concept of the “inscription device,” an instrument that converts material into figures or diagrams for use in scientific work. He notes that once the final inscription is produced, the intermediate stages are often overlooked. In discussing vision and cognition, Latour (2012a) highlights the “immutable mobility” of inscriptions, emphasizing that scientific images emerge from a chain of observations, translations, and communication.
Similarly, this dynamic interplay between machine vision and human vision is observable within the AI laboratory. Here, the employment of images manifests an “experimental system” (Parikka, 2023), wherein the three aforementioned image typologies (as data input/output, as visualization, and as interface) are deployed simultaneously by scientists. Further, by conducting a granular analysis of these data modalities and visualization processes, we can better understand how this interaction across both sides of the interface makes sense. This necessitates a close examination of the situated practices of researchers within the laboratory. As the analysis unfolds, we will see how reintroducing the perspective of the “observer” reveals alternative epistemological commitments within CV.
Operational image and invisibility problem
To understand such human–machine interactions of CV laboratory image practices, I draw on the concept of operational image. It originates in Farocki's analysis of Gulf War targeting footage, where images ceased to function as representational surfaces and instead became components of technical systems that “do not represent an object, but rather are part of an operation” (Farocki, 2004: 13–22). Farocki's foundational insight on operation, that certain images are designed neither for contemplation nor human spectatorship, established a conceptual binary between representational images for humans and operative images for machines. Parikka (2023: xi-19) significantly expands this genealogy, arguing that operational images are not confined to military or industrial contexts but constitute a broad media-technical regime spanning scientific photography, planetary sensing, platform infrastructures, and machine-learning pipelines. For Parikka, the image is best understood through a tripartite ecology of platform, dataset, and model, each indexing a distinct but interlocking layer in how images circulate as computational events (Parikka, 2023: 19–20). In foregrounding this extended field, Parikka shifts the concept from a media-critical observation to a general theory of the invisual architectures of contemporary computation.
A central contribution of the operational image literature is its theorization of invisuality, a condition in which images exceed human optics and become substrates for algorithmic operations (Parikka, 2023: 57–95). Building on Mackenzie and Munster's (2019) notion of “platform seeing”, images today operate primarily as statistical, distributed, and nonoptical events that lack any unified vantage point accessible to human perception. Within platform infrastructures, the individual image is stripped of semantic uniqueness and subsumed into aggregations (ensembles, datasets, and feature vectors) whose value lies in their combinatory potential for training models. This invisual regime is characterized by continual translations between visibility and invisibility: sensor data becomes vectorized; interface images lure users while concealing backend operations; and algorithmic filtering determines which objects become operationally “real.” Parikka (2023: 22) emphasizes that the image's passage from visible surface to invisual data is not merely a technical transformation but a political one, as platforms decide what is captured, counted, or erased. Instead of asking what images mean, scholars such as Hoel (2018) argue that we must ask what images do within heterogeneous assemblages of human and nonhuman agency, allowing media theory to analyze empirical contexts where images oscillate between visible artifacts and invisible computations, revealing the infrastructural processes that render the world machine-readable (Paglen, 2019).
The invisibility problem emerges, as vision is often regarded as a selective perception that reduces information to an operational level, whether individually or collectively (Latour, 2012b; Lynch, 1985). Visibility inherently defines boundaries of invisibility, with the overlooked “invisible” elements frequently becoming critical subjects of study. In a phenomenological style, Merleau-Ponty (1964), Ihde (2001), and Gibson (2014) highlight how tools and perception integrate the body and mind, revealing invisibility as both unperceived phenomena and the cognitive processes underlying vision. Lynch's study of visual documentation reveals that invisibility arises not from reducing visible features but from a sequence of reproductions that progressively modify an object's visibility for abstraction, treating images as both qualitative and quantitative data (Lynch, 1988).
Therefore, the operational image becomes a response to invisibility problem. For example, graphical interfaces now serve as essential units of analysis, foregrounding the operational aspect of scientific knowledge production. Scholars have extended visual analysis to include gestures (Vertesi, 2012), filters (Vertesi, 2014), databases (Jaton, 2017), and learning environments (Passi and Jackson, 2018), exploring how images and meanings evolve dynamically through operation in fields such as protein crystallography (Myers, 2008), radiology (Rystedt et al., 2011), and magnetic resonance imaging (Alač, 2008; Prasad, 2005). These cases demonstrate how selection and mathematization occur across screens, intertwining the visible and invisible through computational visual communication. Operational image calls for a shift in focus—from the belief that images make objects or conclusions visible to examining the moments when images become operational thus invisible. As this paper argues, this shift is rooted in our own expectations of artificial intelligence: we want it to function as a perfectly transparent box, fully visible from every angle, while simultaneously hoping it will display a kind of extraordinary and inherently uncertain creativity.
Invisibility of CV
Among the many scientific practices involving image data and visual communication, CV laboratories provide a concrete and extreme context for analyzing the problem of invisibility. Compared to other research fields, CV exhibits the following distinctive characteristics in handling images: (1) image data serve as the primary and direct research object; (2) tasks include extensive image generation and understanding; and (3) experiments are rooted in computer science, relying on mathematical models and ground truth computation, diverging from the natural sciences paradigm (Morris, 2003). In essence, researchers in CV are the best data scientists at processing images. Their mission is to design a “machine vision” system, enabling computers to “see” (Klette, 2014) by employing a series of techniques to convert images into symbolic descriptions and reconstruct captured scenes from digital images (Horn, 1986).
The divergence between machine vision and human vision amplifies the invisibility problem of image data. As shown above, the process of extending human vision through machines has created more problems of invisibility. As Dobson (2023) notes, machine vision is grounded in human perception, relying on preexisting human knowledge and experience. Gaboury (2021) implies invisibility as a defining characteristic of CV theory. Through an analysis of the “hidden surface problem” in computer graphics—where algorithms resolve invisibility problem—he argues that computer graphics are not composed of visible logic but constructed through processes of data reduction and erasure to align machine vision with human vision. Here, visibility becomes an algorithmic process of withstanding.
In the field of critical CV studies, research emphasizes the dual identity of CV scientists as both social actors and technical workers. This analysis operates on two levels: inwardly, it examines how social identities (Bechmann, 2017; Borji, 2012), positions (Scheuerman and Brubaker, 2024), and biases (Hanley et al., 2021) influence CV algorithms; outwardly, it explores how data scientists negotiate with broader societal stakeholders (Cambo and Gergle, 2022; Hsiao and Shorey, 2023; Passi and Jackson, 2018), thereby shaping the meaning of image data. Scholars further highlight the materiality of technology, engaging with the social construction of “human-centered computer vision system” through audits of artifacts (Prabhu and Birhane, 2020), development (Scheuerman et al., 2021), and data annotation (Miceli et al., 2020).
While upstream data labeling and downstream societal applications have been well studied in the CV workflow, the process of knowledge production within CV remains a significant “black box.” One key challenge lies in the difficulty of documenting the interactions between scientists, computers, and data in existing empirical studies. Therefore, this study adopts an ethnographic methodology rooted in the STS tradition, focusing on the vibrant, hands-on experiences of scientists working with image data in CV laboratories. I adopt invisibility as an epistemological lens to examine the emergence of uncertainty in CV knowledge, its impact on scientific practices and algorithmic design, and the ways in which scientists respond to this epistemic condition. The research questions are proposed as follows:
Methodology and positionality
I conducted a 9-month fieldwork at Lab A, 3 a CV-focused research laboratory situated within a university in Singapore. Lab A is structured into two primary groups based on their tasks: generation and understanding. These groups are further divided into smaller subgroups according to specific projects. Given the independent and nonoverlapping workflows of these groups, I selected one project for in-depth observation.
From December 2023 to May 2024, I conducted observations of the daily activities of Team B in the lab. This period coincided with the development of a research project and the creation of an academic paper from its inception to submission. My observations covered all aspects of the research process, including topic selection, group meetings, data processing, computational experiments, paper writing, and submission. From June to October 2024, I continued following the researchers, conducting monthly visits to discuss their reflections on the project and their rebuttal strategies during peer review sessions, culminating in the paper's acceptance at a top-tier international AI conference. With researchers accessing servers remotely via personal computers from various locations, much of the team's work took place outside the lab. As a result, my observations extended beyond the physical lab to include dormitories, study areas, conference rooms, and even subway commutes.
The research project undertaken by Team B focused on perspective shifting in videos—specifically, generating first-person perspective videos from third-person perspective material. The team described the “novelty” of their work as an original task, largely inspired by a dataset developed within the lab and their prior experience with video-generation algorithms. However, the emphasis on innovation also highlighted the self-referentiality of the project, revealing a black box that this study seeks to unpack.
My specific work included observing the research activities of lab members, auditing group meetings, and conducting in-depth interviews with team members. Complementing these methods, I analyzed the project's associated papers, source code, presentation documents, treating these as components of an “inscription strategy.” To capture the tension between machine and human vision, I integrated my fieldwork with a direct examination of the project's open-source codebase, thereby operationalizing code reading as a methodological tool (Marino, 2020). My analysis centered on how researchers manipulate the code of large models, specifically through parameter freezing and fine-tuning. I explored the translation among text, images, video, and code within the IDE to investigate what is (in)visible to researchers throughout the model training process.
To address the three research questions outlined above, I further refined the research design in Table 2.
Research design of the laboratory ethnography.
Findings
Flashing data
When I first entered the lab, my informant J pointed to the monitors in the computer room and said, “Look, the images appear and disappear… they’re flashing.” Adopting this emic descriptor, I conceptualize the “flash” as a temporal extension of the operational image, a phenomenon characterized by the intermittent visibility of image data during model training. To grasp the model's fundamental task, one might envision a simplified pipeline wherein images exist at both the input and output terminals: determining which visual data feeds into the machine learning process and what visual forms are produced to satisfy human ocular needs. However, the machine learning process itself operates partially autonomously, devoid of direct researcher intervention or real-time parameter adjustments. Consequently, images within the model manifest a basic “flash” structure of “visibility–invisibility–visibility.” Here, a “flash” represents not merely a momentary appearance but the trace left by image computation on the “external retina” (Lynch, 1988), signaling the oscillation between human-readable representation and machine-readable operation.
What makes image data flash? Drawing from the pipeline for this project, I identify three technical affordances contributing to invisibility: data type conversion, data accuracy limitations, and intermodel coupling. These factors replace the optical coherence between the human eye and the research object with computational processes, making direct observation of the image data impossible and creating a visual-level black box.
First, data type conversions occur frequently in the code of models. In the project's open-source code, various data types—RGB images, grayscale images, depth images, and multispectral images—are employed for tasks such as classification, detection, and 3D modeling. Annotated data, including bounding boxes and segmentation masks, also play a critical role in supervised learning and model evaluation. As shown in Figure 1, all four functional modules of the project involve data type conversions. The Exo2Ego view translation prior converts images into features and noise, synthesizes the first frame of the first-person view using scene reconstruction, and feeds it into subsequent models for video generation. Further, two multiview encoder networks (Model A, B) process different feature values to transform images into dynamic (frame-by-frame) and continuous (pose-stabilized) video. Throughout the pipeline, the encode–decode logic abstracts images into computable features, often transforming them into noise or other nonvisual data type. Even on the machine learning monitoring interface, only abstracted loss values 4 are displayed, rendering the raw data unobservable to human eyes.

The pipeline of model design.
Second, data accuracy was frequently sacrificed for computational efficiency. This trade-off produced images so blurred they were unrecognizable to the human eye—a result of both dataset limitations and hardware constraints. At one group meeting, team members calculated that a single Stable Video Diffusion model required 8GB of memory; running two needed 16GB, adding a NeRF model took up another 1–2GB, and intermediate variables pushed the total beyond 70GB. Yet the team's maximum allocation from Singapore's National Supercomputing Center was just 40GB—far short of the pipeline's 80GB + demand. To cope, W spent a week learning DeepSpeed algorithm, attempting to compress image data from 16-bit and BF16-bit floats to 8-bit. The outcome was dismal: “The 16-bit floats were obsolete—just black images. The BF16s 5 weren’t black but turned into blurry mosaics. Totally untrainable.” These digital artifacts are visual symptoms of how invisibility is coded into computational constraints.
Third, model coupling further hinders the researcher's ability to observe images in real time. Cross-computations between different models make it challenging to disassemble computational processes independently, as small algorithmic black boxes combine to form larger, more complex ones. This coupling also significantly increases computational costs: larger models demand more graphics cards, often requiring days of computation and introducing persistent engineering issues (bugs). As a result, observing images becomes increasingly difficult, with the computational process further distancing the researcher's visual perception from the image data. In a conversation, J highlighted the difficulties:
Ideally, it would be simple [laughs]. But most of the time, the systems are deeply intertwined. For example, the outputs from the first layer of Model X's network might need to feed into Model Y, while the second layer of Model X might require adjustments based on the first layer's data before being sent to Model Y. Just stitching models together in a simple way is more of an engineering task—it wouldn’t have much academic value for publication.
The coupling of models amplifies issues related to data types and accuracy, and together, these factors prevent researchers from directly observing the image data. At times, the image data remain hidden within the computational process; at other times, they surface as input or output, shaping the researcher's judgment of the training process. In this dynamic HCI, the image data become “flashing”—serving as the “interface” through which researchers perceive and manipulate the model. Researchers often assume that when the image data are not visible, the machine is performing computations. However, this invisibility creates persistent anxiety, as it signifies a lack of control over the data. To regain visibility, researchers must undertake a series of technical tasks: understanding and “freezing” large models, unifying data types and interfaces, managing memory and data accuracy, and continuously monitoring and validating machine learning performance. These efforts aim to render the data “visible,” allowing researchers to visually perceive and control the experimental flow. Conversely, invisibility suggests that the experiment has spiraled out of control.
Here, flashing data are a critical extension of operational image. It not only elucidates the operability of image data but also theorizes the oscillation of images between visibility and invisibility as a productive and temporal process. While Parikka (2023) provides a foundational critique of operational images, he primarily focuses on the political–economic consequences of invisible images from the perspective of reception. His work offers less detail on the emergence of this invisibility at the site of production: specifically, to whom is the image invisible, at what moments, and under what conditions? As the empirical evidence demonstrates, the “flashing” affordances of image data cause them to recede into the invisible algorithmic black box, while the agency of scientists driven by the institutional imperatives of scientific knowledge production continually demands that these images be rendered visible for interpretation and validation. In the following ethnography, we see laboratory workflows, debugging practices, and metric evaluations reveal that, fundamentally, the crafting of CV is constituted by iterative operational flows of “visibility–invisibility–visibility.” In this sense, the “flash” is the temporal sequence marking the operational image's transformation between states of visibility and certainty. The ethnographic description of the “flash” captures this specific temporal nuance, offering a phenomenological account of how algorithmic uncertainty is managed through visual practice.
Seeking images, sensing models
Unlike traditional scientific laboratories, the CV lab lacks test tubes and microscopes, as well as the pursuit of natural truths. Instead, researchers aim to construct the physics of a cyber world, relying primarily on their own vision as their most critical tool. Figure 2 illustrates the training workflow for the CV model developed by Team B. After inputting the image dataset, the model undergoes multiple iterations (V1, V2, V3, etc.), ultimately passing validation and being inscribed into a finalized form. Model training resembles “building blocks,” where the goal is to determine the appropriate components and their configurations. This process requires experimental design to test different model structures across key stages: data manipulation during iterations, quantitative and qualitative validations, and supplementary testing during peer review and rebuttals.
Although these validations seem to adhere to strict and objective procedures, they are often shadowed by the “ghosts” of flashing images, highlighting the constant presence of human vision. Figure 2 uses different colors to indicate the visibility of images throughout the experimental workflow. Yellow marks represent intermediate states—where image data outputs have been sampled, pretrained, or mixed with other data types—and are therefore only “partially visible.” By correlating the properties of image data with the experimental workflow, this section shows how invisibility shapes the experimental process.

Simplified workflow of model training in the computer vision lab.
Experiment: Bug or failure?
Spend an afternoon in Lab A, and you’ll likely find most researchers frowning at the black background interface of VS Code, tweaking code repeatedly—a clear sign they have run into another bug. In computer science terms, a bug refers to a system malfunction, often considered a straightforward engineering problem. But for scientists training new models, a bug can signal a failed experiment, potentially invalidating their research hypothesis and raising doubts about the feasibility of the entire study. In this context, identifying a bug is not just a technical exercise but a critical process of diagnosing and resolving experimental failures. For Team B, interpreting experimental results through images is the key. When asked how they recognize experimental failures, team members shared a thoughtful response:
This is a tough question to answer. If identifying issues were simple, research wouldn’t be so challenging… I usually have an expected outcome—not necessarily perfect, but if the result is far off, it's likely a bug. For example, if the output image is black and white or completely blank, that is definitely a bug. If it looks terrible, with only faint (shadowy) figures, it's still probably an issue. It's about looking, estimating, and getting a sense of it.
Analysis of experimental workflow.
Broadly, researchers attribute bugs to four potential causes: (1) coding errors, (2) flawed research design, (3) insufficient computational resources, or (4) low-quality datasets. However, pinpointing the exact source of the problem is often elusive. For instance, while running the PixNeRF model, the team repeatedly debugged the code but failed to resolve the poor experimental performance. Ultimately, they set aside the dispute and pushed forward with a new experimental workflow. Debugging, as J shared, occupies more than 70% of his research time. The invisibility of images makes it nearly impossible to discern whether the failure stems from technical faults in the code or fundamental issues with the research design.
J navigates between two visual systems. One belongs to the digital world, comprising monitors, control interfaces, and function curves. The other exists in the physical world, containing experimental controls represented as ground truth—the tangible objects being modeled. Once the images become visible, researchers use their vision to validate the quality of the generated images. Thus, J's interpretations of bugs are deeply tied to vision, with his decision-making occurring after the “flashing” transition from invisibility to visibility. Due to the frequency of this flashing, such judgments are made repeatedly. Focusing on the role of human vision, we might describe the experimental process differently: researchers place photos into a black box, occasionally take them out for inspection, and each time they do, they must evaluate the quality of the images while imagining what is happening inside the box. In essence, the control of flashing data dictates both the pace of the experiment and the direction of the next model iteration.
Validation: Metric or sense?
J's oscillation between two worlds raises a critical question: how is a model validated in the CV lab? How does one assess the quality of an image? According to Team B members, they simultaneously monitor the curves on the control interface and examine the output images, evaluating whether the results align with their “common sense.” Shifting focus from the micro-level experimental processes to the broader context of formal knowledge validation reveals a dynamic interplay between two systems: quantitative and qualitative validation.
The development of quantitative validation techniques is a core research area in machine learning. It underpins the legitimacy of machine learning knowledge and represents the dominant understanding of its logic, computation, and rationality. The loss value, as mentioned earlier by J, measures the difference or error between a model's predicted and actual outcomes. However, the behind-the-scenes workflow of the lab tells a different story of knowledge validation. When the team iterated Model V2 to the next version, they unexpectedly found that adjusting data accuracy led to better visual results. This improvement was not based on quantitative evaluations, such as loss values or metrics, but rather on qualitative validation informed by the researchers’ subjective visual experience and common sense. W elaborated:
I think CV often lacks rigorous evaluation metrics, which makes us overly reliant on human judgment. For instance, when measuring the similarity between two images, we might not have a perfectly accurate method. Sure, there are metrics, but they are not always reliable. It's common to see cases where the metrics look great, but the actual visual result feels wrong.

Comparison of quantitative and qualitative validation results for the model.
This oscillation between quantitative and qualitative validation aligns with Lynch's (1988) discussion of scientific images. Team B's approach to image data does not treat selection and mathematization as independent operations. Selection does not simplify or trim evidence but instead underscores the active role of the researcher's visual judgment. In the generative construction of CV knowledge, visibility is not “gradually modified” into theorization but the model must remain “visible,” with quantitative and qualitative logic presented side by side in the paper to validate the experiment's legitimacy. This uneasy relationship between empirical scientific norms—which strive to exclude subjectivity—and the implicit reliance on human perception creates a persistent tension. As Cetina's (1999) laboratory studies suggest, a purely sensory-free experimental process is a myth; “visual and experiential cues are deeply embedded in the experimental workflow.”
Translating between two worlds through visual communication
The invisibility of image data has significantly reshaped experimental workflows and validation mechanisms. How, then, do AI experts address these impacts? As Cetina and Amann (1990) observe, lab images are not captured but constructed and designed, reflecting science as a communicative process. Building on this, I examine three types of nondata images—sketches, surveys, and renderings—which, while absent from the dataset, mediate researchers’ understanding and demonstrate their agency in making image data “visible.”
Figure 4 shows a sketch drawn by J during data cleaning. It aimed to verify whether the real-world orientations of cameras matched those assumed by the model. Cameras record external parameters, such as position and rotation, but these don’t always align with the model's expectations. On the left side of the sketch, the Z-axes of the spatial coordinate systems for cameras A, B, C, and D are shown pointing in the direction of human vision. The right side illustrates that the cameras’ shooting directions align with the human visual field. According to J's notes, this process aimed to “coordinate” the real world with the machine world because, as J explained, “the model's code assumes one set of camera orientations, while the camera parameters provide another. They don’t always match.”

Hand-drawn sketch by a team member for coordinating data cleaning.
The top-right corner of the sketch specifies the criterion for alignment: the Z-axis must “direct to people.” Establishing this alignment allows real-world scenes to be captured by cameras and subsequently computed by machines, provided both align with human visual orientation. The verification process is straightforward: draw it on paper and make it visible. M noted that while fluency in using sketches and visual experience comes with practice, this step cannot easily be replaced by debugging or other technical methods:
Especially when there are many cameras and parameters, it's easy to get confused. Even reversing one axis can ruin everything. Sometimes, the only way is to draw it out like this. If I handle the parameters incorrectly and just run the code, the results might not even appear. By that point, it's too late to troubleshoot.
Another strategy emerges in response to external critique. During rebuttals and presentations, an intriguing semiotic pattern emerges: reviewers, peer experts, and the lay public—anticipated readers of the paper—are more likely to ask, “Why can’t I see the image here?” rather than, “How does your model explain this?” I argue that questioning the invisibility of the model is more practical than demanding its explainability. Responses to questions about neural network weights or layer design are typically the same: they result from better datasets and computational resources. A tacit norm in the field prefers showing results visually over theoretical explanations.
This is why nearly every section of Team B's rebuttal letters and presentation slides featured illustrative images. These visuals, used to showcase research outcomes, differ from single experimental outputs—they are curated and enhanced representations of results. They appear on GitHub repositories, demo websites, and conference slides, marking the final step in the lab's inscription process. In some cases, these visuals are turned into public surveys where audiences rate them, further reinforcing the reliability of qualitative validation (Figure 5).

Visual survey: Rendering images (left) and scoring forms (right).
Ironically, the scientists who create these visuals also evaluate them. Inscription technologies inevitably involve selective presentation of experimental results, and interpreting these inscriptions becomes both an assessment of experimental validity and a reproduction of the researcher's visual judgment. Lab members offered varied explanations for their use and evaluation of visuals:
Conclusion
An actor–network has come into focus. Anchored in vision, cameras, GPUs, monitors, Team B, scientific communities, and the public all find their place within a CV paper. Scientists oscillate between the two ends of image data, translating between machine vision and human vision. Through an analysis of image data affordances, researchers’ decision-making processes, and various forms of inscription, I observe that the relationships between different actors correspond to specific graphic interfaces. This dynamic and micro-level perspective allows for a closer examination of big data and AI's knowledge production through the lens of operational image.
This study begins with image data, proposing that they gain agency within the network through invisibility. I define the intermittent visibility of images in the lab as flashing data, shaped by three technical affordances: data type conversion, data accuracy limitations, and intermodel coupling. These traits compel researchers to translate image data between different actors. The invisibility of image data introduces uncertainty into the computational process, shifting it from a passive object manipulated and interpreted by researchers to an elusive and operational entity within the model training workflow. Unlike traditional natural sciences, where the experimental process often follows an “entity-data-meaning” knowledge confirmation model (Cetina, 1999), CV operates differently. Its research object is data itself. Yet this study reveals that datasets cannot directly fix meaning; researchers must employ a variety of strategies to reinforce visual understanding, forming a “data-entity-meaning” knowledge confirmation framework.
I further explore how operationality and invisibility challenge conventional scientific experimentation by focusing on two key contradictions: first, the ambiguity between bugs and failures during the experimental process, and second, the divergence between quantitative and qualitative validation standards. The principle of “seeing is believing” persists in machine learning; without visible data, neither technical debugging nor the falsification of knowledge is feasible. This positions CV models as a trading zone for meaning-making. The invisibility of images fractures the traditional “hypothesis–observation–validation” chain into two segments: “hypothesis–invisibility” and “invisibility–validation.” Observation, instead of being a discrete step, permeates other stages of the process. This explains why experiment disputes are sometimes set aside, model structure adjustments appear arbitrary, and general quantitative metrics occasionally lead to counterintuitive conclusions.
Thus, developing new visions to address the invisibility of raw image data becomes crucial. This study categorizes researchers’ strategies into two types: self-enhanced vision and external vision incorporation, each addressing the contradictions outlined above. The former guides researchers in troubleshooting and redesigning model structures, while the latter brings social vision into the collective negotiation of models. Double cleaning or reinterpretation of data balances professional expertise with common sense, facilitating the continuous renegotiation of power both within and beyond the laboratory. In this way, the claim that “give me a laboratory and I will raise the world” gains new significance. Once scientific beliefs achieve broad societal acceptance, the laboratory's magic no longer depends solely on specialized instrument spaces. Instead, it must be sufficiently normalized and socialized to ensure that the perceptions of ordinary people are validated before it can acquire the authority to represent the world.
A processual analysis of crafting machine vision introduces an interactive perspective that reinterprets the epistemological foundations of artificial intelligence. Nearly every stage of machine vision training requires vision involvement, signifying that machine learning operates through dynamic HCI. In CV laboratories, knowledge production relies on visual infrastructures embedded with researchers’ perceptual structures. This compels us to view artificial intelligence as a highly sensory technoscience, a technology-in-use, and even an artistic practice. Such a perspective challenges the notion of AI as nonempirical, posthuman, and reliant on aggressive induction driven by immense computational power. Moreover, the anxiety surrounding the opacity of machine learning aligns with humanity's incomplete understanding of its own visual communication mechanisms. The body—the computer's operator—emerges from the vast dataset, while researchers navigate numerous graphical interfaces. It is their interactions with hardware that drive ever-larger-scale computations. “Humans and machines must operate simultaneously,” J shared as the secret of his research. While the machine rests, his eyes remain vigilant.
Footnotes
Acknowledgements
The author is deeply grateful to the anonymous lab members who generously “co-coded” this research with me. The author also extends sincere thanks to the reviewers and the editorial team for helping finalize this manuscript. Special thanks to Xiuli Wang, Daniele Macuglia, Lili Lai, and Tso Kwok for their insightful feedback and continuous encouragement. This article builds upon my master's thesis and was awarded the Top Student Paper Prize by the AEJMC (2025) Visual Communication Division, where audience feedback helped refine my arguments. Finally, the author’s heartfelt thanks to the Peking University Alishan Scholarship for enabling my exchange in Singapore, where the idea of this paper was truly born.
Informed consent statement
Access to the laboratory and participation in this study were granted with informed consent from all participants. To protect the privacy and confidentiality of the interviewees, certain details have been anonymized in the presentation of the findings.
Funding
The author received no financial support for the research, authorship, and/or publication of this article.
Declaration of conflicting interests
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
