Abstract
Architectural design shapes cultural identity, particularly in regions like the UAE, where architecture blends heritage and modernity. However, generative AI struggles to accurately capture architectural heritage due to biases in training data and model interpretability. This study evaluates six generative models—Flux, PromeAI, Stable Diffusion XL, DeepAI, Getimg.ai, and ImagineArt—in producing designs across three categories: traditional, modern, and hybrid (modern inspired by traditional). Using CLIP scores to measure textual-image alignment, we found traditional styles were most faithfully generated, while modern designs proved more challenging. Hybrid styles achieved moderate performance, reflecting the complexity of merging cultural motifs. Flux and PromeAI outperformed others, excelling in traditional and hybrid categories. These results underscore the potential and limitations of AI in architectural design, emphasizing the need for refined datasets and training techniques to improve cultural accuracy. By integrating AI tools, architects can enhance conceptual design efficiency, experiment with cultural elements, and streamline ideation. Future research should prioritize culturally diverse datasets, multimodal learning (e.g., 3D generative approaches), and real-time AI-assisted workflows. This study advances AI-driven architectural innovation while promoting culturally authentic design practices.
Introduction
Artificial intelligence (AI) is being used in architecture in the form of generative design tools, virtual assistants, and machine learning algorithms. 1 Generative design tools generate design alternatives based on user-supplied constraints and objectives. Virtual assistants help architects manage deadlines and workload, freeing up time for creative tasks. Machine learning algorithms can improve building performance by learning from historical building data. AI-powered construction tasks and monitoring systems could increase efficiency and reduce human labor. AI-powered sensors and monitoring systems could also predict maintenance and repairs, allowing building owners to schedule tasks in advance. 2
The rapid advancements in AI, particularly in generative tools, have significantly influenced creative fields, from art to architecture. Through sophisticated machine learning tools, text-to-image generative AI tools can convert textual prompts into incredibly detailed visuals, allowing designers to visualize ideas. Recent research on generative AI in architectural design has categorized its applications into six key areas: concept image generation, architectural 3D form generation, plan generation, facade generation, and structural system generation. 3 These advancements highlight the growing role of AI in shaping architectural workflows, from the early ideation stages to detailed structural planning. However, while generative tools can produce innovative and visually compelling results, their effectiveness depends on the quality of the training data and their ability to integrate cultural and contextual nuances, which represent essential considerations when applying AI to architectural heritage preservation and contemporary design in culturally significant regions. 4
AI-driven generative tools have demonstrated remarkable capabilities in producing realistic architectural designs by manipulating image data based on predefined parameters. These tools, particularly Generative Adversarial Networks (GANs) 5 and Deep Neural Networks, 6 synthesize new artworks by blending various artistic styles, often mimicking the aesthetics of specific designers. GANs are widely employed to generate images from scratch, making the diversity and quality of their training datasets crucial to the architectural design process. In the context of text-to-image tools, the generated visuals are influenced by the prompts provided, as these tools retrieve and reinterpret images based on textual descriptions and associated signifiers. 7
Preserving a city’s history and culture is vital for future generations, as it helps sustain the city’s identity. Cultural and symbolic elements play a key role in bringing people together and fostering a sense of security. 8 AI’s ability to blend traditional knowledge with new technology will be crucial in maintaining the various cultural landscapes that characterize humanity’s shared heritage as it develops further. By integrating AI-driven generative tools into architectural workflows, 9 designers can not only enhance creative exploration but also ensure that cultural heritage remains a central element in future urban developments. This capability has particular significance in regions like the United Arab Emirates (UAE), where architectural practices are deeply rooted in cultural heritage yet push the boundaries of modernity. The UAE’s architectural landscape is a convergence of traditional Islamic and Emirati influences with contemporary global styles, 10 resulting in structures that embody both cultural identity and innovation. Text-to-image generative tools offer a unique means to explore this blend, providing architects with a dynamic kit to create designs that reflect Emirati culture while embracing modern trends.
This study investigates the potential of AI-generated outputs in representing aesthetic and creative cultural heritage identity, specifically analyzing outputs from leading text-to-image AI tools to assess their alignment with local UAE heritage. It explores how AI-generated images can enhance the incorporation of traditional Emirati elements in modern architectural practices. These elements include islamic geometric patterns, mashrabiya screens, and the use of courtyards. Exploring how AI tools can help architects create culturally resonant designs will provide insights on the role of AI in cultural preservation within architectural design.
By utilizing these tools, designers can experiment with cultural motifs in ways that were previously limited by conventional design software, paving the way for a more fluid and interactive design process. This work builds on ongoing discussions regarding AI’s applications in architecture, particularly its capacity for supporting cultural continuity and innovation. Despite advancements in generative AI, a critical research gap persists concerning the accurate representation of non-Western architectural styles, particularly those rooted in UAE’s rich cultural heritage. Existing tools are primarily trained on datasets that mostly feature Western architectural elements, resulting in biased or culturally inaccurate outputs when applied to Eastern design contexts. 11 The objective of this study is to address the challenge of dataset bias in existing generative AI tools by systematically evaluating how they can be refined as support mechanisms for ideation and conceptualization in architectural design, thus better capturing the intricacies of the UAE’s architectural identity, contributing to more culturally authentic AI-driven design processes as a result.
The remainder of this paper is organized as follows: The next section presents the Background and Literature Review, introducing the fundamental concepts related to generative AI and architectural design, in addition to summarizing prior research on AI-assisted design and its implications. It consists of three subsections: Generative AI Technology Evolution, which outlines the history of generative AI development, Generative Tools, which discusses various AI-driven generative tools, and finally, Architectural Design Process, which outlines the conventional design workflow and the integration of AI tools. Afterwards, the Research Methodology section describes the experimental setup and evaluation metrics used in this study. The subsequent section presents the Results and Discussion, analyzing key findings and their significance in relation to existing research. Finally, the Conclusion section provides a summary of insights, limitations, and directions for Future Work.
Background and literature review
An outline of the development of generative AI technologies and their increasing incorporation into workflows for architectural design is provided in this section. Understanding the underlying tools and their capabilities becomes essential as the field of architecture evolves to embrace AI, especially in the creative and conceptual phases.
Generative AI technology evolution
There has been significant progress in text to image generative AI tools, utilizing different structures to produce high-quality and lifelike images. Generative Adversarial Networks (GANs), like StyleGAN and BigGAN, utilize a generator-discriminator structure to produce visually striking images by improving results through adversarial training.12–15 Diffusion tools such as DALL·E 2 and Stable Diffusion utilize textual prompts to guide the iterative transformation of random noise into detailed images, showcasing excellence in photorealism and fine details.16,17 Transformer-based tools like Contrastive Language-Image Pre-training guided diffusion utilize extensive language-image pairings to gain a profound understanding of text prompts, allowing for consistent image creation that matches the descriptions. 18 Hybrid approaches blend these methods, increasing control, accuracy, and variety in results. GANs excel in speed, diffusion tools offer precision, and transformers bring robust understanding, making them essential in various applications, from design to entertainment, as each type has its strengths. 19
Generative AI has come a long way, with each breakthrough pushing the boundaries of what artificial intelligence can do. From the early days of rule-based systems like ELIZA in the 1960s
20
to the cutting-edge tools we have today, such as GPT-4 and diffusion-based technologies,
21
the progress has been remarkable.
4
Figure 1 shows how AI has evolved from simple neural networks to sophisticated systems capable of creating human-like text, images, and more.
22
This journey has unlocked new possibilities, transforming fields like art, design, and countless other areas with creative and practical applications. The following timeline traces AI’s journey from theoretical beginnings to transformative tools shaping industries today: • 1960s – ELIZA: ELIZA, an early chatbot developed by Joseph Weizenbaum, marks one of the first examples of generative AI. It simulated human-like conversation but was rule-based rather than learning-based.
20
• 1980s–1990s, Neural Networks: The development and use of neural networks became prominent, enabling foundational work in deep learning and AI. This era introduced architectures like Recurrent Neural Networks (RNNs) and Convolutional Neural Networks (CNNs).
23
• Early 2000s, Deep Learning: Deep learning techniques, powered by more complex neural networks and GPUs, started to revolutionize AI, enabling breakthroughs in fields like image recognition and text processing.
24
• 2014 - GANs: GANs, introduced by Ian Goodfellow, were a major leap in generative AI. They enabled the creation of realistic images and other types of media by pitting a generator network against a discriminator network.
13
• 2015, Diffusion Tools: Diffusion tools emerged as a new paradigm for generating high-quality images by iteratively denoising data. These tools gained widespread attention with tools like Stable Diffusion and DALL-E.
25
• 2020, GPT-3: OpenAI’s GPT-3 represented a breakthrough in language tooling, demonstrating how large-scale transformers could generate human-like text. Its technology laid the foundation for multimodal tools like DALL-E.
26
• 2023, Bard and Watsonx: Recent advancements, including Google’s Bard and IBM’s Watsonx, highlight the latest efforts to refine AI systems for text generation, dialogue, and multimodal content creation.
27
Key events in the history of generative AI.

Generative tools
Text-to-image tools have revolutionized the way we generate visual content, leveraging advancements in AI to create detailed, photorealistic, and conceptually rich images from textual descriptions. DALL·E, 17 developed by OpenAI, uses advanced diffusion techniques to generate creative and highly detailed artwork, making it a leader in the field. MidJourney 28 is another prominent tool that focuses on producing artistic and stylistically unique images, popular among creative professionals. Diffusion tools, 25 as a foundational technology, have become a key framework for generating high-quality images, with iterative noise removal and refinement. GPT, 26 though primarily a language tool, is increasingly integrated with visual tools to generate image descriptions or support multimodal systems. Tools like PromeAI 29 and DeepAI 30 make AI art generation accessible to a broad audience, offering versatile tools for image creation. Flux 31 brings cutting-edge advancements in diffusion tools, while Imagine Art and Getimg.ai 32 provide specialized features for generating and customizing images, catering to specific creative or professional needs. Together, these tools empower users to bring their ideas to life through seamless AI-driven creativity. These tools have been selected in this study based on their relevance to the focus on culturally grounded architectural image generation. Each represents a distinct approach to the image generation workflow, from diffusion tools to prompt engineering methods, which allows for the evaluation of how different techniques visualize culturally specific architectural prompts.
Architectural design process
The architectural design process goes through several stages. As observed in Figure 2, the first step is to define the problem and analyze the program to better understand the functional requirements, usage patterns, budget, client perspective, and other factors. Problem identification involves conducting a Strengths, Weaknesses, Opportunities, and Threats (SWOT) analysis, organizing information through mapping, gathering inspirations, and brainstorming.
33
The second step consists of information gathering, which is done with the aim of defining the scope of the design project, providing a better understanding of the different factors involved.
34
The third step involves ideation and schematic outline for the purpose of addressing the problem that was identified in the first step.
35
The fourth step of the architectural design process is the design development stage, which transforms conceptual ideas into detailed design outputs. It involves 2D and 3D visualizations, physical tools, and detailed drawings that communicate the design solutions for iterative refinement.
36
Presentation and communication are carried out in the fifth step, in which the proposed designs are communicated to stakeholders through tools, drawings, and physical or digital tools to accommodate their point of view and feedback.
37
The sixth step involves detailed design and documentation, producing technical papers and specifications for further guidance, after which construction is carried out in the final step to translate the created design into practical reality.
38
Traditional architectural design process.
25

The integration of generative AI into the architectural design process can significantly help architect conceptualize, visualize, and communicate their design ideas. As traditional architectural design relies on time-consuming approaches like manual sketching and physical tools, more complex design problems require rapid iteration. 28 This motivates the integration of AI tools in design, which can enhance the efficiency and creativity in all stages of the design process. Generative AI in particular offers designers new methods of ideation and visual exploration for different cultural and aesthetic needs. 3 Comparative studies are crucial to understanding which generative AI tools are most effective in specific architectural design applications. Recent research has extensively explored the potential of text-to-image generative AI tools, such as MidJourney, DALL-E, and Stable Diffusion, particularly in Stage 3 of the architectural design process, where initial ideas are developed.28,39,40 This stage is pivotal for exploring creative directions and refining concepts that later form the foundation for Step 4 (Design Development). The comparative study presented in this work examines how different text-to-image generative AI tools perform in supporting the conceptualization and schematic design stage. As creative exploration is critical in this stage, a comparative evaluation of these tools helps identify the most effective systems for transforming textual architectural prompts into culturally relevant and coherent visual outputs. The presented comparative framework in this study contributes to identifying the potential roles and limitations of generative AI tools for the architectural design process, particularly in culturally sensitive non-western regions.
The ideation stage involves brainstorming, exploring creative possibilities, and generating preliminary ideas which shape the conceptual direction of a design. This phase focuses on understanding user needs, spatial functions, and aesthetic visions, where designers seek to experiment with broad conceptual frameworks. Studies have shown that generative AI tools like MidJourney and Stable Diffusion have been instrumental in enhancing this phase by enabling rapid generation of diverse design and novel concepts.28,41 An important factor in the effectiveness of these tools is the input textual prompts, whose clarity and context directly influence the quality and relevance of the generated outputs. 42 While these tools foster creativity, they often reflect inherent biases from their training datasets. For example, Al-Haroun 7 found that generative tools struggled to authentically replicate Gulf architectural styles, often defaulting to generic, Western-inspired motifs that lacked cultural specificity. 43 The study highlighted that these tools excelled at replicating globally recognized landmarks but misrepresented less-documented cultural elements, such as Islamic geometric patterns and Emirati wind towers (barjeel). This suggests that AI-generated ideation may inadvertently skew toward Western-centric aesthetics, thereby limiting the cultural diversity and authenticity of early design concepts. 44 This limitation becomes even more evident in the conceptualization stage, where abstract ideas are transformed into more concrete architectural and visual forms. 11
The conceptualization stage refines preliminary ideas into tangible concepts by integrating design principles, technical constraints, and project requirements. Designers translate abstract notions into schematic sketches, tools, or visualizations, forming the foundation for the design development phase. The quality and accuracy of the input prompts also play an important role in this stage, as they highly determine how well the generated outputs reflect the designer’s intended forms, materials, and cultural details. 45 In this stage, research by Zhang et al. 46 demonstrated that while AI tools can generate visually appealing outputs, they struggle with cultural coherence and architectural accuracy when conceptualizing non-Western designs. For instance, AI-generated outputs intended to reflect Islamic architecture often lacked the intricate details that define authentic designs, such as accurate renditions of calligraphy, geometric facades, and traditional motifs. Moreover, Bandi et al. 47 emphasized that the absence of culturally diverse datasets contributes to these limitations. Without region-specific data, AI tools tend to homogenize architectural concepts, leading to oversimplified or culturally inaccurate outputs. This challenge is particularly critical in conceptualization, where accuracy and attention to cultural details are important.
The following section reviews the recent advancements in generative AI tools for architecture from 2021 to 2024, categorized to align with key areas of architectural design and AI integration. The first category, evaluation studies of generative tools in architecture, is divided into two subsections: generative tools for ideation, focusing on their role in the initial creative process, and generative tools for conceptualization, examining their application in refining design concepts based on the Architectural Design Process in the background section. The second category addresses comparative studies of generative tools in architecture, analyzing their performance, usability, and impact on design workflows. The third category explores methodologies and techniques in generative AI for architecture, highlighting the technical foundations, tools, and innovative approaches employed. Finally, challenges, limitations, and research gaps in AI-driven architecture are discussed, shedding light on unresolved issues and opportunities for future research.
Evaluation studies of generative tools in architecture
This section discusses research that evaluated the contributions of generative AI tools to various phases of the architectural design process. The studies are further divided into two primary areas to give a structured overview: ideation and conceptualization, which summarize recent research under each of these themes.
Generative tools for ideation
Paananen et al. 39 explored how these tools support creative processes in architecture by providing visual stimuli and enhancing ideation workflows through a laboratory study with architectural designers. Their research highlighted that these tools could be pivotal in early design stages, enabling architects to generate multiple iterations rapidly, although challenges were noted in achieving architectural accuracy and coherence. Similarly, Albaghajati et al. 48 focused on how designers utilize these tools for concept generation, finding that text-to-image tools could enhance creativity, especially in the early stages of the design process. However, they are not yet precise enough for technical architectural designs.
The study by Yildirim et al. 2 explores the integration of text-to-image AI in architecture, focusing on tools like DALL-E 2, MidJourney, DiffusionBee, 49 and MotionLeap. 50 It explores the history of AI, its relationship with architecture, and its applications in architecture, interiors, and urban design. These tools enhance creative ideation by allowing architects to translate textual descriptions into images. While promising in ideation and visualization, challenges such as reliance on extensive datasets, lack of technical precision, and potential misinterpretation of textual prompts limit their utility in detailed architectural design. The paper also highlights the future potential of transitioning to 3D tool generation and immersive design. Another study 51 examines the usability of text-to-image generative AI tools in architectural design education. A workshop was conducted with architecture students who explored these AI tools to generate architectural designs. The results showed that while the tools were helpful for ideation and early design stages, they were less effective as the designs became more complex, especially when generating technical architectural details like plans and sections.
The work by Ploennigs et al. 52 explores the applicability of diffusion-based AI art tools in architectural design, particularly in ideation and visualization. It evaluates how MidJourney, DALL-E 2, and Stable Diffusion handle various architectural tasks such as sketching, style transformation, and detailed plans. The authors collected around 85 million MidJourney user queries over a period of 1 year to analyze the usage trends of AI for architectural designs. They filtered the user queries for terms related to architectural design, like “architect”, “interior”, and “exterior”. They found 5.7 million queries with architectural intent, with 2.2 million explicitly specifying design terms. The authors proposed workflows that combine the strengths of these different tools to address architectural use cases based on their findings.
Generative tools for conceptualization
Hakimshafaei 53 surveyed Midjourney, DALL-E, and Stable Diffusion, evaluating their respective capabilities in producing architecture-specific images through experiments and comparisons. It demonstrates how AI tools generate complex forms and designs that are challenging for human designers. The research highlights the efficiency of AI tools in producing varied design options quickly compared to traditional algorithmic tools. However, the quality and adherence to architectural principles varied. AI tools were found to be better suited for early-stage ideation than for detailed, technical work. Another work 54 examines how Midjourney supports the brainstorming process for generating the conceptual form of a Safavid mosque. The study evaluates Midjourney’s ability to create designs with creativity, speed, and adherence to traditional mosque proportions, which is considered a representative example of Safavid mosque architecture. The findings indicate that Midjourney is reliable in terms of speed and creativity, but it lacks accuracy and adherence to inputs.
With further details, Horvath and Pouliou 55 explored the application of AI tools like text-to-text, text-to-image, and image-to-image generation for conceptual architectural design, specifically for the eVolo skyscraper competition. Through a research-through-design approach, the authors examine how these AI tools can assist in creating new design briefs and exploring different shapes, and inspiring innovative design solutions in architecture. While some generated texts lacked clarity, they still offered valuable insights for conceptual development. Anna Jaruga-Rozdolska 28 explores the use of MidJourney in architectural design by testing its ability to produce visualizations based on text inputs for specific themes such as underwater restaurants or Baroque-style façades. The study evaluates how these tools can assist architects in the early conceptual stages, and their outputs were compared against human-generated sketches in conceptual quality and adaptability. While the tool generates aesthetically pleasing and conceptually relevant images, it cannot deliver complete architectural designs, relying heavily on user intervention and prompt engineering for further development.56,57
Comparative studies of generative tools in architecture
Ploennigs et al. 58 explored how generative AI tools (like ChatGPT and Midjourney) interact with architectural history. They examined whether these tools can accurately distinguish between different architectural styles or if they generate incorrect (hallucinated) information. The study includes both qualitative and quantitative analyses of more than 101 million Midjourney queries to identify trends in the use of AI for generating architectural images. They found that ChatGPT demonstrated a broad understanding of architectural styles but often hallucinated information, particularly regarding specific architects and examples, whereas Midjourney generally produced accurate representations for well-documented styles but struggled with certain non-Western or less-known styles. A similar study 1 explored the role of AI, specifically generative algorithms and search engines, in transforming architectural design. It evaluated how AI algorithms help designers in creative exploration, highlighting the potential and limitations of tools like DALL-e, mid-journey, and other AI-based systems (e.g., ChatGPT) in the design process. AI’s ability to process vast datasets, generate images, and refine design concepts suggests a shift in the way architects and designers engage with tools for ideation and visualization.
Another study 46 compared designs produced by Antoni Gaudí with those generated by an AI diffusion tool. Participants evaluated them across five key dimensions: authenticity, attractiveness, creativity, harmony, and overall preference. Gaudí’s designs were rated higher in authenticity and harmony, while AI designs showed potential, particularly in attractiveness and creativity. The study highlights both the possibilities and challenges of using AI in architectural design and emphasizes the need for improvement in AI’s ability to capture the human elements of design. Similarly, Yousef Al-Haroun 7 explores the influence of AI on architectural design and cultural identity in the Gulf region, focusing on Kuwait, Riyadh, and Doha. Through workshops with architecture students, AI was employed to reinterpret significant local landmarks. Findings revealed that AI excels in replicating globally recognized structures but often misrepresents local cultural elements, suggesting digital bias.
Methodologies and techniques in generative AI for architecture
Methodological advances in the application of generative AI for architecture are essential for improving tool performance and relevance. Li et al. 3 presented a systematic literature review on techniques used in text-to-image generation for architectural design, highlighting methods for enhancing image realism and architectural accuracy through tool adaptation. Their study emphasized the use of customized datasets and tool fine-tuning to bridge the gap between architectural standards and the output of generative tools. One prominent example in the literature of fine-tuning techniques is LoRA (Low-Rank Adaptation), 59 which is a technique used to efficiently fine-tune large pretrained tools like Stable Diffusion. Instead of updating all the tool’s weights (which is computationally expensive), LoRA injects small trainable low-rank matrices into certain layers. These matrices learn task-specific adjustments, allowing the tool to adapt to new datasets or styles without retraining the entire tool. This approach can help fine-tune models for specific architectural styles, as done by Zhang et al. 60 for instance. There are several other studies in the literature which focus on integrating generative AI into design tasks. Shi et al. 61 proposed an AI-powered approach for automating exterior architectural conceptual design. The proposed system combines textual and non-verbal design intents, like client needs, architectural language, and sketches, to generate architectural images using a fine-tuned variant of Stable Diffusion. The ControlNet tool, which enables the usage of extra inputs like sketches and depth maps in text-to-image tools, 62 was used to guide the generation process by reflecting the design intent expressed in sketches. Validation against two existing tools demonstrated that this method can produce design alternatives based on the provided architectural design intent. Similarly, Ma and Zheng 16 proposed a method for generating building facades using Stable Diffusion. The tool was trained using the low-rank adaptation (LoRA) method, which injects small trainable low-rank matrices into the layers of the tool instead of updating all of it during training to reduce time and computational constraints. 59 The training data was acquired from the Center for Machine Perception (CMP) Facades dataset, which includes 606 façade images from different sources, to improve the accuracy of generated architectural images. 63 ControlNet is used to enable more precise control over the generated outputs, addressing the challenge of randomness in text-to-image generation. The study demonstrates that this method can efficiently generate diverse building facades based on textual prompts with controllable style. Jummung Chen et al. 64 extend this approach by training a diffusion tool on a newly constructed dataset of architectural designs. Their tool generates designs with specific styles and high visual quality, improving design efficiency and quality in the conceptual design phase and helping reduce reliance on manual rendering efforts.
Generative AI has also demonstrated significant potential in the cultural heritage context. Garozzo et al. 12 proposed a hybrid method to enhance the understanding of cultural heritage scenes by combining knowledge-based and data-driven approaches. Their method uses GANs to create realistic images for training image classification and retrieval systems. The GANs are anchored to a semantic ontology domain, which can generate both isolated objects and full scenes that aid in automating the classification and retrieval of cultural heritage images. Another study 65 explored the use of text-to-image tools to digitally reconstruct heritage sites that have been damaged or destroyed by conflicts and natural disasters. By generating images from textual descriptions informed by historical records, the AI tools aim to recreate accurate visual representations of the original structures. This innovative approach bridges the gap between traditional architectural restoration and modern digital reconstruction, offering significant potential for cultural heritage preservation.
Challenges, limitations, and research gaps in AI-driven architecture
The advancements in generative AI integration with architectural design have led to many tools and approaches, as observed in the reviewed literature. Figure 3 illustrates the overall distribution of explored tools in architectural research, where diffusion-based tools (30%), MidJourney (28%), and DALL·E (20%) dominate the field, demonstrating a notable preference for text-to-image diffusion tools. Table 1 further categorizes these studies, showing that most research focuses on evaluation, comparison, and methodological development, with relatively limited attention to cultural and regional applications. This pattern indicates that although architectural designers have experimented with various tools, like GANs and VAEs, they continue to prioritize tools that output high-fidelity designs with efficiency. Despite these advances, key challenges and limitations remain. Bandi et al.
47
reviews the foundations of generative AI, identifying issues like the absence of standardized evaluation metrics, the challenge of achieving diversity and realism in design contexts, and the high computational requirements and ethical concerns regarding data bias and authorship. Similarly, Enjellina et al.
40
acknowledge the advantages of generative AI in the form of accelerating the conceptualization process and expanding creative exploration. However, they also emphasize that the effective use of these tools depends heavily on user prompt engineering and tool adaptability to architectural contexts. More specifically, Sukkar et al.
41
explored MidJourney’s capability to represent Islamic architectural heritage. They identify key limitations that include linguistic and cultural biases in prompt interpretation, inconsistencies in representing regional architectural styles, and inaccuracies in fine details like ornamentation and calligraphy. Their study supports the need for more inclusive and culturally informed generative AI training datasets and workflows. While generative AI tools have been explored with interest in architectural research, the presented analysis of tool distribution and usage patterns uncovers an imbalance in the form of a high reliance on general-purpose AI tools with limited domain-specific and culturally aware adaptation. Distribution of tools explored in the literature review. Tools explored in each paper in the literature review.
The literature review reveals a significant gap in recent studies on generative AI for architecture. While substantial research has been conducted on prominent tools such as MidJourney, DALL·E, and diffusion-based systems, emerging tools like PromeAI, Flux, Imagine Art, and Getimg.ai remain underexplored. A key observation, based on our review of the existing literature, finds that most studies focus on performance and creativity in Western design contexts or general architectural aesthetics rather than culturally authentic applications. Despite the increasing popularity of these tools, there remains a lack of systematic evaluation benchmarks for cultural relevance of the generated designs. This can be significant for regions affected by rapid modernization like the UAE, where architectural heritage is being increasingly overshadowed by global design trends.10,66,67 As most of the analyzed generative AI tools tend to reproduce Western-centric designs, there is notable a risk of overlooking regional identity elements like wind towers and adobe textures. Centering our study within the UAE context enables testing the adaptability of these emerging AI tools to culturally specific architectural contexts and contributes to research on the localization of AI for diverse cultural heritages.
Methodology
The integration of generative AI tools into architectural design has opened new avenues for creativity and innovation, enabling designers to explore complex forms, cultural motifs, and stylistic interpretations with unprecedented ease. Text-to-image tools have demonstrated remarkable potential in generating context-aware visuals that align with specific design goals. For regions like the UAE, where architectural identity is deeply rooted in cultural heritage yet shaped by modern innovation, these tools offer unique opportunities to blend tradition with contemporary aesthetics.
This research focuses on systematically evaluating the capabilities of leading text-to-image AI tools in generating culturally significant architectural designs tailored to the UAE context. The study examines their performance in rendering intricate details of traditional architecture, such as Islamic geometric patterns, wind towers, and courtyards, alongside modern design elements that reflect the UAE’s progressive architectural ethos.
Figure 4 shows the methodology diagram that outlines a systematic approach for generating and evaluating UAE-specific cultural heritage images using generative AI tools. It begins with Finding a Gap, identifying unexplored areas in integrating generative AI with architectural design, particularly for UAE’s cultural heritage. Next, Model Selection involves choosing appropriate unexplored AI tools capable of producing high-quality, culturally relevant images. Building Prompts follows, where detailed descriptions are crafted to guide the tools. This step loops with Testing Prompts, where the generated outputs are assessed, and adjustments are made in cooperation with professional architect to refine the descriptions. The loop continues until Final Prompts are achieved, representing the most effective input for generating the desired images. Using these finalized prompts, Image Generation is performed to create high-quality visuals. The outputs are then subjected to Image Evaluation, assessing their relevance, quality, and alignment with the intended cultural and architectural themes. Finally, Results Interpretation involves analyzing the generated images to draw meaningful insights and conclusions about the tool’s capabilities and applications in the chosen domain. Methodology.
Tools
Identifying appropriate tools and tools based on their technical capabilities and output quality is the first step in evaluating the generative capacity of AI tools in architectural design, particularly in capturing the details of the UAE’s cultural heritage. The second step involved analyzing how various prompt structures affect the quality and cultural alignment of the generated imagery. The methodical assessment of certain tools and the creation of various prompt forms that direct the generation process are described in depth in the following subsections.
Exploration of different tools
The first step in this study involved exploring various text-to-image tools and tools to identify the most suitable ones for evaluation. Given the rapid advancements in generative AI, numerous tools have been developed, each with unique capabilities, strengths, and limitations. This exploration phase focused on understanding the differences in tool architectures, training datasets, and output quality to ensure a comprehensive evaluation. By analyzing multiple tools, the study aimed to select tools that best align with the research objectives, particularly in generating culturally relevant architectural designs.
Given the diversity of available tools, the exploration focused on understanding their underlying architectures, training methodologies, and output quality. The tools tested included Stable Diffusion XL 68 and DALL·E 3, 17 both accessed through the Phygital + web interface, as well as DALL·E Mini via Craiyon. Additionally, VQGAN + CLIP 69 was tested using Colab code, while Flux 70 was evaluated through its dedicated web interface. This comparative analysis provided insights into each tool’s strengths, limitations, and ability to generate culturally relevant architectural designs.
To complement this evaluation, various text-to-image tools were also explored. These included PromeAI, which utilizes Controllable AI-Generated Content (C-AIGC) tools, Imagine.art, which employs custom diffusion tools, Getimg.ai, built on Stable Diffusion, and DeepAI, which leverages a neural network-based approach. Each tool was assessed based on its usability, output quality, and suitability for generating culturally relevant architectural imagery. This comparative analysis provided a comprehensive understanding of how different AI tools and tools contribute to text-to-image generation in architectural design.
Exploration of different prompts structure
To assess the ability of these tools and tools to generate Gulf, Arabic, and Islamic heritage-inspired architectural designs, a structured set of prompts was developed across four levels of complexity. The Simple level consisted of brief prompts (2–3 words) focusing on fundamental architectural elements. The Medium level (4–6 words) provided slightly more descriptive inputs to refine the generated outputs. The Detailed level (7–10 words) incorporated specific architectural features, styles, and cultural elements. Finally, the Full Description level consisted of comprehensive prompts with intricate details, including materials and contextual aspects (see Appendix A, Table 5 for the prompts). This multi-level approach allows for a systematic evaluation of how well each tool captures and represents cultural heritage in architectural design generation.
To ensure consistency and conduct a more in-depth exploration of the tools’ ability to capture aesthetic and architectural aspects, a second trial was conducted using a new set of 10 prompts (see Appendix A, Table 6 for the prompts). These prompts followed a structured format to provide more detailed and contextually rich inputs: [main subject/location], [details of subject/materials], [background], [art/image style], [lighting], [colors], and [keywords]. This approach allowed for a more comprehensive assessment of how well each tool interpreted intricate architectural elements, cultural motifs, and visual aesthetics. By incorporating specific attributes such as materials, lighting conditions, and artistic styles, the trial aimed to evaluate the tools’ precision in generating high-quality, culturally accurate architectural imagery.
The selection of these attributes was based on their influence on how generative tools interpret and create architectural concepts. For instance, materials like adobe or coral stone indicate vernacular authenticity within Arabian and Islamic architecture. Likewise, specifying styles or aesthetic qualities helps evaluate how these tools balance creative interpretation with culturally authentic representation. The selection of these elements thus aims to standardize the level of detail across different prompts for fair comparison between the tools and test the tools’ capacity to interpret and reproduce culturally significant visual elements. This approach connects prompt design with the study’s objective of understanding the effectiveness of generative AI tools in producing culturally accurate architectural designs in the UAE context.
Prompting
Prompts used for each category.
To ensure unbiased testing of generative AI tools, the location and building names are intentionally omitted when generating the images. This approach allows the tools to focus solely on architectural and stylistic features without direct contextual anchoring, enabling an objective evaluation of their ability to replicate culturally and architecturally accurate representations.
Image generation
Six prominent text-to-image tools were selected to perform a systematic evaluation in this study. PromeAI, Imagine.art, Getimg.ai, DeepAI, Stable Diffusion XL, and Flux were selected based on criteria specific to architectural design research. The selection process considered (1) technical performance (image quality, realism, and level of detail), (2) creativity and stylistic flexibility, (3) cultural relevance, in terms of the ability to express UAE architectural motifs, (4) user-friendliness and accessibility for students, and (5) suitability for iterative design workflows. These factors are linked to the study’s objective of assessing the capability of generative AI to support architectural designers, especially in the UAE’s cultural design context.
Every tool was evaluated according to specific factors that are crucial for architectural design. A comparison of the tools by tool type, interface accessibility, and prompt requirements can be observed in Table 3. • PromeAI: Known for its advanced generative capabilities, particularly in producing highly detailed and context-aware images. This aligns with the need to create delicate, culturally relevant designs that showcase regional materials, shapes, and motifs in the context of UAE architectural heritage. Despite its usefulness, this tool can be difficult for exploring architectural design because of its sharp learning curve and dependence on prompt engineering, emphasizing the balance between advanced capabilities and user accessibility. • Imagine.art: Focuses on stylistic and artistic interpretations, producing results that are highly abstract and creative. Because of this, it is especially helpful during the architectural exploration phase of conceptual design, when stylistic experimentation is crucial. Its adaptability in expressing both traditional and modern concepts makes it a valuable tool for blending tradition with modern design, which is the study’s main goal. • Getimg.ai: Designed for high-quality image generation with a focus on usability. Its straightforward interface and capability to handle detailed and realistic outputs make it well-suited for producing architectural designs that require a balance between creativity and feasibility. This’s ability to generate varied outputs from a single prompt facilitates iterative design processes, aligning well with traditional studio-based learning methods. • DeepAI: Excels in abstract and creative image generation, enabling users to explore unconventional and innovative designs. Its simplicity and user-friendly interface make it accessible to a broader audience, including students and educators. In the context of architectural design, its versatility in producing both abstract and detailed outputs supports a wide range of creative and practical exploration relevant to UAE heritage and modernity. • Stable Diffusion XL: Recognized for producing high-resolution, photorealistic images with advanced control over the output. For this study, it serves as an excellent benchmark for creating detailed architectural renderings that reflect the intricate patterns and textures characteristic of UAE cultural heritage. Its ability to incorporate specific design elements, coupled with its customization capabilities, ensures outputs that align with the study’s emphasis on cultural authenticity. • Flux: Released by StabilityAI, Flux brings cutting-edge generative capabilities with a focus on speed and efficiency. Developed with a blend of German engineering and AI expertise, it provides a unique lens for generating images with a balance of precision and creativity. Its scalability and ability to handle complex prompts make it particularly useful for exploring intricate architectural forms and cultural motifs, essential for UAE-focused architectural designs. Comparison of AI tools by their features.
Evaluation
To generate images for the selected generative AI tools, a standardized approach is employed to ensure consistency and comparability across outputs. All images are generated using a 1:1 aspect ratio, a realistic visual style, and a single image per prompt. These settings are chosen to maintain uniformity, reduce variability, and focus on the tools’ ability to produce high-quality, culturally accurate representations of UAE landmarks.
As for evaluating the generated images, CLIP (Contrastive Language–Image Pretraining) is a machine learning model that aligns text and images in a shared embedding space, allowing it to assess how well an image corresponds to a given text description.18,71,72 The similarity score is computed based on cosine similarity between their respective embeddings. 73 Higher CLIP scores indicate a closer alignment, reflecting the tool’s ability to capture the architectural details and stylistic elements described in the prompt. Using the CLIP score ensures an objective and quantitative evaluation of the generative AI tools, helping to identify which tools produce the most accurate and contextually relevant images.
Figure 5 provides an overview of how CLIP works. It illustrates the architecture and training process of CLIP, which is designed to learn a joint representation of images and text. Here’s a breakdown of the key components and steps in the figure: 1. Contrastive Pretraining • Input: CLIP takes a batch of image-text pairs as input. For example, an image of a dog is paired with a caption like “Pepper the Aussie pup.” • Encoders: ○ Text Encoder: The input text is tokenized and processed by a transformer-based text encoder, generating a fixed-dimensional embedding for the text. ○ Image Encoder: The input image is processed by a convolutional neural network (CNN) to produce a corresponding fixed-dimensional embedding for the image. • Embedding Space Alignment: Both the image and text embeddings are mapped into a shared latent space. The tool learns to maximize the similarity between embeddings of matching image-text pairs while minimizing the similarity between mismatched pairs. This is achieved using a contrastive loss function. • Embedding Matrix: The embeddings for all text and image pairs in a batch are compared pairwise. This forms a similarity matrix where each element represents the similarity score between a specific image and text embedding. 2. Zero-Shot Classification: • Dataset Preparation: For classification tasks, labels are transformed into text prompts (e.g., “A photo of a {object}” such as “A photo of a dog”). • Text Encoding: The text prompts are encoded using the text encoder, resulting in embeddings for each class label. • Image Encoding: Images are encoded using the image encoder to generate their embeddings. • Similarity Matching: The similarity between the image embedding and each class label’s embedding is computed. The class label with the highest similarity score is predicted as the class for the image. 3. Zero-Shot Prediction: • The pre-trained CLIP tool can generalize to new tasks without additional fine-tuning. For example, given an unseen image and a set of descriptive text prompts, CLIP can identify which prompt best matches the image. • Joint Representation Learning: By learning a joint embedding space, CLIP aligns visual and textual information, enabling robust cross-modal retrieval and classification tasks. • Contrastive Loss: The tool optimizes a contrastive loss to ensure that semantically related image-text pairs are close in the embedding space, while unrelated pairs are far apart. Visual summary of the CLIP mechanism.
71

This approach allows CLIP to excel in zero-shot learning tasks, where it can predict labels or perform retrieval tasks for data it hasn’t explicitly been trained on. The embedding-based approach ensures a flexible and generalizable framework for linking images and text.
Results and discussion
The CLIP score analysis offers valuable insights into the performance of the evaluated tools across three architectural categories: traditional, modern, and modern inspired by traditional. Below is a detailed interpretation and comparison of the results shown in Table 5. The results of the additional experiments are presented in Appendix C in Tables 7–9.
Performance across categories
The average CLIP score for every tool in each category is summarized in Figure 6. The results indicate that the traditional category achieved the highest average CLIP scores across most tools, highlighting the tools’ capability to reproduce cultural and architectural heritage accurately. For instance, Flux scored the highest in this category with an average CLIP score of 0.355, followed closely by deepAI at 0.3484 and PromeAI at 0.3446. In contrast, the Modern category generally yielded lower average scores, suggesting that contemporary designs pose more challenges for the tools. Modern designs inspired by traditional designs exhibited moderate performance, reflecting the complexity of blending modernity with traditional elements. Average performance for each category in each tool.
Performance across prompts
Analysis across individual prompts revealed variations in performance. Traditional prompts, detailed in Table 2, consistently received higher CLIP scores, especially prompt 5, which achieved high scores across most tools. This suggests that prompts emphasizing well-defined traditional elements are more effectively interpreted by the tools. For modern prompts, the performance was relatively lower, with prompt 3 showing marginally better results. For the modern Inspired by traditional category, prompts that included a mix of traditional elements yielded balanced yet moderate scores, highlighting room for improvement in handling hybrid designs, as observed in Figure 7. CLIP scores for each category, evaluated using the corresponding prompts listed in Table 2.
Tool-specific observations
The performance of each generative tool was evaluated across three architectural categories: Traditional, Modern, and Modern inspired by Traditional. Figure 8 shows each tool and its scores across the categories for each prompt, while Figure 9 compares the scores of each tool within the same prompt category. The images generated using each tool across the different architectural categories can be observed in Table 4 alongside their CLIP scores. DeepAI demonstrated consistent performance across all categories, excelling particularly in the Traditional category with a score of 0.3484, while showing a slight decline in Modern (0.3098). GetimgAI exhibited the weakest overall performance, with the lowest average scores across all categories. Its highest score was in the Traditional category (0.3304), slightly surpassing its performance in Modern inspired by traditional (0.3). ImagineArt maintained a steady performance across categories, achieving a notable average score of 0.3332 in the Traditional category and moderate results in Modern and Modern Inspired by Traditional. PromeAI showed strong results in the Traditional category (0.3446) while delivering moderate performance in the remaining categories, particularly excelling in prompts that blended modern and traditional elements. SDXL displayed a balanced performance, with its highest score in Traditional (0.332), slightly outperforming other categories, while Modern inspired by traditional scored 0.3138 on average. Among all tools, Flux delivered relatively better results, with a slightly stronger performance in Traditional prompts (0.355) and maintaining competitive scores across other categories, indicating more robustness compared to other tools in diverse scenario. However, the differences in the scores remain incremental and demonstrate the competitiveness and close performance of all the analyzed tools. As can be observed in Figure 9, most tools demonstrated marginally better performance in the Traditional prompts category. CLIP scores for each tool across different categories and corresponding prompts. CLIP Scores by prompt type for each tool. Clip scores across traditional, modern, and modern inspired by traditional.

General trends and insights
Several trends emerged from the analysis. The tested tools generally excelled in generating images using the Traditional prompts in Table 2, indicating a strong alignment with well-documented and structured cultural features. The Modern category posed challenges, potentially due to the abstract and diverse nature of contemporary designs. Hybrid prompts (Modern inspired by traditional) performed moderately, revealing a gap in the tools’ ability to seamlessly blend traditional and modern elements. The prompt-specific analysis underscored the importance of clear, detailed, and culturally grounded descriptions in achieving higher CLIP scores. Additionally, Flux consistently outperformed other tools, suggesting its robustness in handling varied architectural and design scenarios. The CLIP score findings in this study were used to indicate conceptual alignment and visual relevance, with the results demonstrating that traditional and hybrid styles performed better due to them being more structured and recognizable at the ideation level.
Implications for future research
An essential factor which should be considered in future research is the presence of bias in training data. Despite the improvement of generative AI for abstract and hybrid design in recent years, mainstream training datasets remain primarily representative of Western architectural styles, which leads to an underrepresentation or misinterpretation of non-Western cultural elements. This bias was demonstrated through the selection of the UAE as a case study, which presented traditional Emirati designs that were difficult for generative tools to authentically replicate or blend with modernity.
Future research should focus on addressing this limitation through several potential methods, like fine-tuning tools with datasets that emphasize modern and hybrid architectural styles while ensuring a balanced and diverse representation of different cultural traditions. Integrating context-aware training could further enhance the tools’ ability to interpret cultural and design nuances, reducing biases that lead to homogenized or inaccurate outputs. Additionally, investigating the potential of combining multiple generative AI tools to leverage their strengths in different categories could mitigate individual tool biases and enhance overall performance.
Developing human-AI collaborative frameworks through standardized benchmarks and metrics tailored for evaluating generative AI performance in architectural and cultural heritage contexts is essential to systematically assess and address biases. Moreover, incorporating user feedback from architects, historians, and designers familiar with regional architectural styles could help refine prompt structures, improve tool interpretability, and enhance usability in real-world applications.
These directions can pave the way for more effective, culturally sensitive generative AI tools that contribute meaningfully to architectural design and heritage preservation. Beyond that, the exploration of real-world applications of generative AI, such as urban planning, digital heritage site reconstruction, and sustainable housing design, also holds significant potential for future research. Generative AI can contribute to other applications like interior design and educational tools which enhance learning in architectural studies.
Conclusion
This research evaluated the performance of multiple generative AI tools in creating architectural designs across three categories: traditional, modern, and modern inspired by traditional. The CLIP scores showed small fluctuations across the three tested prompt categories. This indicated that all the analyzed tools performed in similar scoring bands regardless of prompt type. In the traditional category, most tools trended between 0.33 and 0.37 with occasional in Prompt 3 and Prompt 4 for specific tools. The modern prompt category showed slightly lower score fluctuation, with scores ranging from around 0.28 to 0.33. This suggests greater consistency, but slightly weaker overall alignment compared to the traditional category. The modern inspired by traditional prompts showed moderate variability and a generally upward movement across tools like promeai, deepai, and Flux. The analyzed patterns indicated that prompt type moderately influences performance, with traditional prompts yielding the highest CLIP alignment while modern prompts gave slightly lower but more stable scores. Hybrid prompts, on the other hand, demonstrated balanced performance with slight improvement in certain tools. These results highlight the importance of targeted improvements in training methodologies, dataset diversity, and prompt optimization to enhance tool performance. Addressing these challenges can pave the way for generative AI to play a more significant role in architectural design, effectively combining creativity with precision to meet diverse stylistic and contextual demands.
It is necessary to stress that the examined generative tools in this study are most effective in early design stages. They function as creative aids and visual reference generators for architectural designers, not final design solutions, and thus cannot replace human expertise, which remains essential for refinement, validation, and technical development. Integrating these tools within the architectural design process presents designers with opportunities to enhance creativity, efficiency, and cultural relevance across the different stages of design. In the conceptual stage, these tools can rapidly visualize numerous design alternatives using textual prompts, which supports early-stage ideation and helps architects explore culturally relevant designs and materials. In the development of schematics, generative AI can assist with the refinement of façade compositions and stylistic coherence, transforming abstract ideas into tangible design outcomes.
The tools evaluated in this study can provide authors with initial visual references that can be further developed with conventional design software. Heritage-inspired projects, such as those situated in the UAE context, can benefit from generative AI tools as they can aid in representing vernacular architectural elements within heritage-inspired concepts. However, careful consideration is required for the meaningful integration of AI. Dataset curation must take into consideration potential biases and misrepresentations to ensure the unbiased training of these AI tools, and the generated outputs must be evaluated not just for visual fidelity but also cultural relevance by domain experts.
Supplemental material
Supplemental Material - Generative AI in architecture: Examining text-to-image models and platforms for cultural heritage representation in UAE design
Supplemental Material for Generative AI in architecture: Examining text-to-image models and platforms for cultural heritage representation in UAE design by Farah Abu Hamad, Iman Ibrahim, Manar Abu Talib and Mohamed Al Hemairy in International Journal of Architectural Computing.
Footnotes
Acknowledgements
The authors would like to express their sincere gratitude to the University of Sharjah and the OpenUAE Research and Development Group for their support of this research project. We also acknowledge the valuable contributions of our research assistants in conducting experiments, implementing code, and assisting with data collection and analysis.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: University of Sharjah and the OpenUAE Research and Development Group.
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Supplemental material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
