Sample theme detail
Sign in to create your own trend reports and explore all themes
LLM Creativity Evaluation and Generation
Theme #2Scoping Review: LLM Creativity Evaluation and Generation
Overview
This research theme examines how to measure and enhance creative capabilities in large language models (LLMs)—AI systems trained on vast amounts of text data. The field addresses a fundamental challenge: while LLMs can generate fluent and coherent text, researchers are working to understand whether this output is truly creative or simply recombination of familiar patterns, and how to evaluate and improve creative performance across diverse domains like writing, science, and advertising.
Research Landscape
-
Creativity Evaluation Frameworks: Researchers are developing comprehensive benchmarks and metrics to assess creativity across multiple dimensions, as seen in CreativityPrism: A Holistic Benchmark for Large Language Model Creativity (arxiv:2510.20091), What Shapes a Creative Machine Mind? Comprehensively Benchmarking Creativity in Foundation Models (arxiv:2510.04009), and Rethinking Creativity Evaluation: A Critical Analysis of Existing Creativity Evaluations (arxiv:2508.05470).
-
Domain-Specific Creativity Assessment: Specialized evaluation approaches are being developed for particular fields including marketing, scientific research, and creative writing, as demonstrated by Creativity Benchmark: A benchmark for marketing creativity for large language models (arxiv:2509.09702), Large Language Models for Scientific Idea Generation: A Creativity-Centered Survey (arxiv:2511.07448), and CLAWS:Creativity detection for LLM-generated solutions using Attention Window of Sections (arxiv:2510.17921).
-
Understanding LLM Creative Limitations: Studies investigate why LLMs tend toward safe, generic outputs rather than truly novel ideas, examining factors like model confidence and cultural bias in Galton's Law of Mediocrity: Why Large Language Models Regress to the Mean and Fail at Creativity in Advertising (arxiv:2509.25767), Confidence, Not Perplexity: A Better Metric for the Creative Era of LLMs (arxiv:2510.08596), and Cultural Alien Sampler: Open-ended art generation balancing originality and coherence (arxiv:2510.20849).
-
Enhancing Creative Generation: Researchers are developing methods and systems to push LLMs toward more innovative outputs through guided exploration, structured workflows, and deliberate techniques, as shown in Magellan: Guided MCTS for Latent Space Exploration and Novelty Generation (arxiv:2510.21341), Algorithm Generation via Creative Ideation (arxiv:2510.03851), and Spacer: Towards Engineered Scientific Inspiration (arxiv:2508.17661).
-
Creative Writing and Narrative Generation: Work focuses on improving LLM performance in storytelling and creative writing through better datasets, process-level understanding, and multi-agent systems, including COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes (arxiv:2510.14763), Style Over Story: A Process-Oriented Study of Authorial Creativity in Large Language Models (arxiv:2510.02025), and CreAgentive: An Agent Workflow Driven Multi-Category Creative Generation Engine (arxiv:2509.26461).
Knowledge Gaps
-
Theoretical Grounding of Creativity Metrics: While multiple evaluation frameworks exist, there is limited consensus on which metrics best capture genuine creativity versus surface-level novelty, and how different metrics relate to human creative judgment across cultures and contexts.
-
Cross-Domain Creativity Consistency: Most research focuses on isolated domains (writing, science, marketing), leaving unclear how creative capabilities transfer across different types of tasks and whether improvements in one domain benefit others.
-
Process-Level Understanding: The field lacks deep investigation into how LLMs generate creative outputs—the internal reasoning and decision-making processes—compared to extensive focus on evaluating final outputs.
-
Personalization and Subjective Creativity: Limited research addresses how to evaluate creativity relative to individual preferences and cultural contexts, as most benchmarks assume universal definitions of what constitutes creative work.
Methodological Approaches
-
Human-Centered Evaluation: Researchers use expert human judges, crowdsourced comparisons, and established psychological creativity tests (like the Torrance Test) to validate computational metrics, as demonstrated in Curiosity-Driven LLM-as-a-judge for Personalized Creative Judgment (arxiv:2510.05135) and A Comparative Approach to Assessing Linguistic Creativity of Large Language Models and Humans (arxiv:2507.12039).
-
Multi-Dimensional Analysis: Rather than single scores, researchers decompose creativity into measurable components (originality, coherence, diversity, validity) and analyze how models perform across these dimensions, as shown in HypoSpace: Evaluating LLM Creativity as Set-Valued Hypothesis Generators under Underdetermination (arxiv:2510.15614) and The Geometry of Creative Variability: How Credal Sets Expose Calibration Gaps in Language Models (arxiv:2509.23088).
-
Guided Search and Structured Generation: Techniques like Monte Carlo tree search, knowledge graphs, and agent-based workflows are employed to steer models away from default patterns and explore more novel conceptual spaces.
Future Directions
-
Unified Creativity Framework: Developing a standardized, theoretically-grounded creativity evaluation system that works across domains and languages could enable more meaningful comparisons between models and clearer progress tracking in the field.
-
Real-World Creative Collaboration: Research should explore how LLMs can effectively collaborate with human creators in practical settings (design, research, marketing), moving beyond benchmarks to understand actual creative value and user satisfaction.
-
Interpretability of Creative Processes: Investigating what makes certain model architectures, training approaches, or prompting strategies more conducive to creativity could reveal actionable insights for improving creative capabilities at scale.