TY - JOUR
T1 - Survey on large model-empowered visualization and visual analytics
AU - Wang, Yunhai
AU - Cao, Nan
AU - Chen, Siming
AU - Li, Chenhui
AU - Zeng, Wei
AU - Tao, Jun
AU - Zeng, Qiong
AU - Wang, Changbo
AU - Zhang, Jiawan
N1 - Publisher Copyright:
© 2026, Editorial and Publishing Board of JIG. All rights reserved.
PY - 2026
Y1 - 2026
N2 - Data visualization, as a fundamental technology supporting human cognition and data-driven scientific reasoning, is undergoing a profound shift driven by the rapid development of large-scale models and is attracting widespread attention from the academia and industry. For decades, traditional visualization approaches have mainly relied on manually or automatically designed visual encoding rules, graphical grammars, and human-computer interaction techniques. Through explicit visual mappings, graphical composition, and interface operations, these methods enable users to observe patterns, understand structures, and communicate insights from data. However, with the continuous growth of data size, task complexity, and decision-making contexts, statistical visual mappings and parameter-driven interaction are increasingly insufficient to support modern analytical needs. Instead, users now demand more than visual presentation alone; they require systems capable of semantic understanding, task-driven reasoning, and cross-modal information integration. The emergence of large models—particularly large language models, vision-language foundation models, and intelligent agents—has introduced unprecedented technological momentum to the visualization field. These models are not only transforming how visualizations are generated and interacted with but also reshaping the theoretical foundations, analytical workflows, and narrative practices of data visualization. Their strong capabilities in semantic representation, abstraction, and reasoning enable visualization systems to move beyond surface-level depiction toward deeper support for analytical intent and knowledge construction. From a theoretical perspective, visualization research has long been grounded in the Grammar of Graphics and declarative specification languages, which provide highly abstract and formalized descriptions of graphical structures, data mappings, and interaction logic. The introduction of large models does not replace this; instead, it reinforces its importance. Declarative grammars increasingly function as an interpretable "intermediate language" between natural-language intent and executable visual representations, enabling large models to translate high-level analytical goals reliably into controlled, consistent visualization specifications. This mediation is critical for ensuring the interpretability, reproducibility, and controllability of automatically generated visualizations. Moreover, the semantic reasoning capabilities of large models allow them to infer the underlying intent and logic behind visual layouts, color encodings, and spatial structures, and pushes visualization research from low-level perceptual encoding toward semantic-driven visual understanding and intent modeling. At the same time, techniques such as differentiable rendering, neural implicit representations, and Gaussian splatting provide new frameworks for high-fidelity scientific rendering, continuous data representation, and optimization in parameterized spaces. These methods enable visualization systems to become differentiable, learnable, and estimizable, and allows tighter integration between visual representation, computational models, and analytical objectives. As a result, visualization is increasingly situated within richer signal and representation spaces, supporting adaptive, expressive, and controllable visual analysis. Beyond theoretical advances, large models are fundamentally changing how users collaborate with visualization systems. Traditionally, effective visualization required users to master visualization grammars, tool-specific operations, and data processing workflows. By contrast, large-model-driven systems allow users to express analytical goals, design preferences, and data-related questions directly through natural language. The system can then interpret user intent, generate appropriate visualizations, restructure data views, or optimize visual encodings accordingly. This fusion of model-based semantic knowledge with formal visualization representations forms the foundation of a new generation of intelligent visualization systems. At the level of visual analytics, large models and agents are driving a transition from human-centered interaction toward a hybrid collaborative paradigm involving humans, agents, and knowledge. Early machine-learning-for-visualization research primarily focused on addressing isolated subtasks, such as view recommendation, feature detection, or layout optimization. By contrast, visualization frameworks based on large language models (LLMs) leverage unified semantic understanding across tasks and modalities to support end-to-end analytical pipelines. These systems can assist with visualization generation, pattern discovery, and reasoning-based explanation, and enable more holistic support for complex analytical processes. As human-machine collaboration evolves from simple question answering toward knowledge generation, large models increasingly function as proactive cognitive collaborators rather than passive assistants. They can produce trend interpretations, support hypothesis testing, and summarize underlying mechanisms while continuously modeling user behavior, analytical state, and task context. This approach enables a shift from reactive assistance to mixed-initiative collaboration, where systems actively guide users through complex, dynamic, and uncertain analytical environments, augmenting human reasoning and sensemaking capabilities. In the domain of narrative visualization, large models mark the beginning of a new era of intelligent content generation. Traditional narrative visualization requires expertise in design, storytelling, writing, and programming, which makes producing high-quality narratives costly and time consuming. Large models significantly lower this barrier by automatically identifying data themes, extracting salient trends, constructing narrative structures, and integrating visual, textual, and multimodal elements into coherent, stylistically consistent narratives. This capability substantially improves efficiency and accessibility and enables end-to-end automation from data to story in applications such as data journalism, educational visualization, science communication, and business reporting. With the maturation of multimodal generation technologies, narrative processes can further incorporate natural-language interaction, dynamically adapt narrative perspectives, and tailor content to audiences with different backgrounds and goals. In visualization evaluation, large models demonstrate promising potential due to their understanding of visual aesthetics, layout quality, perceptual principles, and readability. They can automatically assess visualization quality, detect misleading encodings, propose design improvements, and generate multiple design alternatives for comparison. However, model-based visual judgment is not inherently equivalent to human perception, a situation raising critical research challenges. Ensuring alignment between model assessments and human visual cognition, mitigating hallucinations and biases, and improving transparency, reliability, and interpretability remain central issues in large-model-driven visualization research. To organize the rapidly evolving landscape of large-model-powered data visualization systematically, this paper presents a comprehensive survey from four perspectives: visualization fundamentals, visual analytics, narrative visualization, and visualization evaluation. This paper reviews representative research advances in each area, analyzes emerging technologies enabled by large models, and discusses key challenges and future research directions. By providing a structured knowledge map and theoretical framework, this survey aims to support future innovation in intelligent visualization systems and contribute to the development of data visualization as a critical bridge between human intelligence and artificial intelligence in the era of large models.
AB - Data visualization, as a fundamental technology supporting human cognition and data-driven scientific reasoning, is undergoing a profound shift driven by the rapid development of large-scale models and is attracting widespread attention from the academia and industry. For decades, traditional visualization approaches have mainly relied on manually or automatically designed visual encoding rules, graphical grammars, and human-computer interaction techniques. Through explicit visual mappings, graphical composition, and interface operations, these methods enable users to observe patterns, understand structures, and communicate insights from data. However, with the continuous growth of data size, task complexity, and decision-making contexts, statistical visual mappings and parameter-driven interaction are increasingly insufficient to support modern analytical needs. Instead, users now demand more than visual presentation alone; they require systems capable of semantic understanding, task-driven reasoning, and cross-modal information integration. The emergence of large models—particularly large language models, vision-language foundation models, and intelligent agents—has introduced unprecedented technological momentum to the visualization field. These models are not only transforming how visualizations are generated and interacted with but also reshaping the theoretical foundations, analytical workflows, and narrative practices of data visualization. Their strong capabilities in semantic representation, abstraction, and reasoning enable visualization systems to move beyond surface-level depiction toward deeper support for analytical intent and knowledge construction. From a theoretical perspective, visualization research has long been grounded in the Grammar of Graphics and declarative specification languages, which provide highly abstract and formalized descriptions of graphical structures, data mappings, and interaction logic. The introduction of large models does not replace this; instead, it reinforces its importance. Declarative grammars increasingly function as an interpretable "intermediate language" between natural-language intent and executable visual representations, enabling large models to translate high-level analytical goals reliably into controlled, consistent visualization specifications. This mediation is critical for ensuring the interpretability, reproducibility, and controllability of automatically generated visualizations. Moreover, the semantic reasoning capabilities of large models allow them to infer the underlying intent and logic behind visual layouts, color encodings, and spatial structures, and pushes visualization research from low-level perceptual encoding toward semantic-driven visual understanding and intent modeling. At the same time, techniques such as differentiable rendering, neural implicit representations, and Gaussian splatting provide new frameworks for high-fidelity scientific rendering, continuous data representation, and optimization in parameterized spaces. These methods enable visualization systems to become differentiable, learnable, and estimizable, and allows tighter integration between visual representation, computational models, and analytical objectives. As a result, visualization is increasingly situated within richer signal and representation spaces, supporting adaptive, expressive, and controllable visual analysis. Beyond theoretical advances, large models are fundamentally changing how users collaborate with visualization systems. Traditionally, effective visualization required users to master visualization grammars, tool-specific operations, and data processing workflows. By contrast, large-model-driven systems allow users to express analytical goals, design preferences, and data-related questions directly through natural language. The system can then interpret user intent, generate appropriate visualizations, restructure data views, or optimize visual encodings accordingly. This fusion of model-based semantic knowledge with formal visualization representations forms the foundation of a new generation of intelligent visualization systems. At the level of visual analytics, large models and agents are driving a transition from human-centered interaction toward a hybrid collaborative paradigm involving humans, agents, and knowledge. Early machine-learning-for-visualization research primarily focused on addressing isolated subtasks, such as view recommendation, feature detection, or layout optimization. By contrast, visualization frameworks based on large language models (LLMs) leverage unified semantic understanding across tasks and modalities to support end-to-end analytical pipelines. These systems can assist with visualization generation, pattern discovery, and reasoning-based explanation, and enable more holistic support for complex analytical processes. As human-machine collaboration evolves from simple question answering toward knowledge generation, large models increasingly function as proactive cognitive collaborators rather than passive assistants. They can produce trend interpretations, support hypothesis testing, and summarize underlying mechanisms while continuously modeling user behavior, analytical state, and task context. This approach enables a shift from reactive assistance to mixed-initiative collaboration, where systems actively guide users through complex, dynamic, and uncertain analytical environments, augmenting human reasoning and sensemaking capabilities. In the domain of narrative visualization, large models mark the beginning of a new era of intelligent content generation. Traditional narrative visualization requires expertise in design, storytelling, writing, and programming, which makes producing high-quality narratives costly and time consuming. Large models significantly lower this barrier by automatically identifying data themes, extracting salient trends, constructing narrative structures, and integrating visual, textual, and multimodal elements into coherent, stylistically consistent narratives. This capability substantially improves efficiency and accessibility and enables end-to-end automation from data to story in applications such as data journalism, educational visualization, science communication, and business reporting. With the maturation of multimodal generation technologies, narrative processes can further incorporate natural-language interaction, dynamically adapt narrative perspectives, and tailor content to audiences with different backgrounds and goals. In visualization evaluation, large models demonstrate promising potential due to their understanding of visual aesthetics, layout quality, perceptual principles, and readability. They can automatically assess visualization quality, detect misleading encodings, propose design improvements, and generate multiple design alternatives for comparison. However, model-based visual judgment is not inherently equivalent to human perception, a situation raising critical research challenges. Ensuring alignment between model assessments and human visual cognition, mitigating hallucinations and biases, and improving transparency, reliability, and interpretability remain central issues in large-model-driven visualization research. To organize the rapidly evolving landscape of large-model-powered data visualization systematically, this paper presents a comprehensive survey from four perspectives: visualization fundamentals, visual analytics, narrative visualization, and visualization evaluation. This paper reviews representative research advances in each area, analyzes emerging technologies enabled by large models, and discusses key challenges and future research directions. By providing a structured knowledge map and theoretical framework, this survey aims to support future innovation in intelligent visualization systems and contribute to the development of data visualization as a critical bridge between human intelligence and artificial intelligence in the era of large models.
KW - agents
KW - fundamental theories of visualization
KW - large language models
KW - visual analytics
KW - visual storytelling
KW - visualization interaction
UR - https://www.scopus.com/pages/publications/105042692542
U2 - 10.11834/jig.260046
DO - 10.11834/jig.260046
M3 - 文章
AN - SCOPUS:105042692542
SN - 1006-8961
VL - 31
SP - 2198
EP - 2221
JO - Journal of Image and Graphics
JF - Journal of Image and Graphics
IS - 6
ER -