ZETTSU LAB

Research

Our laboratory studies AI orchestration technologies that coordinate multiple AI models, databases, external services, computing resources, and human knowledge according to the purpose at hand.
AI orchestration selects and connects appropriate AI models and data in response to a given problem or real-world situation, while continuously improving workflows and models through evaluation of their results.
By combining LLMs, multimodal AI, predictive models, databases, external evaluators, edge AI, and other components, we aim to realize recognition, reasoning, decision-making, and service delivery that are difficult for a single AI model to achieve.

Conceptual illustration of AI model and agent collaboration (Image generated with ChatGPT)

AI Model and Agent Collaboration

We coordinate multiple AI systems with different areas of expertise to achieve recognition, reasoning, generation, and decision-making that are difficult for a single AI model.
External tools and evaluators are also used to verify results and, when necessary, trigger regeneration or replanning.

View Research Topics

Conceptual illustration of knowledge and data orchestration (Image generated with ChatGPT)

Knowledge and Data Orchestration

We collect and organize video, audio, text, sensor, and spatiotemporal data, and store them as knowledge that AI systems can use.
By searching and referring to relevant information, AI systems can reason and generate outputs based on real-world conditions and past cases.

View Research Topics

Conceptual illustration of a distributed and continuously evolving AI platform (Image generated with ChatGPT)

Distributed and Cyclically Evolving AI Platforms

We securely and efficiently coordinate data and computing resources distributed across clouds, edge devices, and organizational environments.
Our goal is to build AI platforms that improve through use by feeding real-world experience and evaluation results back into model updates.

View Research Topics

AI Model and Agent Orchestration

Multiple AI models and agents with complementary capabilities are coordinated to support recognition, reasoning, generation, and decision-making beyond the capabilities of a single model. The approach also integrates databases, external tools, and evaluators to verify generated results and revise plans when necessary.

Research Topics

  • Safe Action Support through the Integration of LLMs and Physical AI
    Action candidates generated by physical AI are evaluated using LLMs and external evaluators, and are revised, replanned, or delegated to humans when necessary. This approach helps prevent unsafe behavior by autonomous vehicles, robots, and other AI systems operating in the physical world.
  • AI Agent Support for Development Process Design
    Ambiguous natural-language requirements are progressively decomposed and structured to support requirements definition, development process design, and content generation. In applications such as game development, multiple AI agents are assigned complementary roles and execution sequences to support iterative collaboration between humans and AI.
  • Advanced Dialogue Understanding and Hypothesis Generation through LLM Collaboration
    LLM-based generation and multidimensional evaluation are combined to extract contextual and causal information from uncertain multi-turn dialogues. Applications include causal hypothesis generation from dialogue logs and response generation that better reflects user intent, thereby improving the understanding and reliability of conversational AI.
  • Multimodal AI Model Orchestration
    Pretrained models for video, audio, text, and sensor data are flexibly combined to build accurate recognition and prediction models. By integrating dashboard-camera footage, heart-rate data, and in-vehicle environmental sensor data, the approach can predict near-miss driving situations that are difficult to detect from video alone.
Multimodal AI model collaboration

Multimodal AI model collaboration that combines pretrained models for different data modalities to create highly accurate multimodal prediction models.

Knowledge and Data Orchestration

Multimodal information, including video, audio, text, dialogue, sensor, and spatiotemporal data, is structured, stored, retrieved, analyzed, and reused as events and human-centered knowledge. By connecting the resulting knowledge databases with LLMs, the approach enhances reasoning, generation, and summarization.

Research Topics

  • Multimodal Video Summarization with Temporal Context
    Events extracted from videos are organized in a knowledge database and retrieved together with preceding and following context and similar examples to generate coherent summaries. The approach can identify important moments in long-form videos, such as esports broadcasts, and produce summaries that reflect the overall flow of a match.
  • Knowledge Structuring of Human Subjective Information
    Human emotions, expectations, reasons for evaluation, and decision rationales are extracted from dialogue and behavioral data as structured knowledge. Applications include modeling dissatisfaction as a gap between expectations and actual outcomes, with the aim of improving customer support and dialogue assistance.
  • Event Data Warehouse
    Sensing data from weather, environmental, traffic, and IoT sources are converted into a common event format and organized by spatial, temporal, and thematic attributes. Useful patterns discovered from large-scale data can support traffic risk analysis, environmental monitoring, and the generation of high-quality training and synthetic data.
  • Spatiotemporal Data Mining
    High-utility patterns are efficiently discovered from event data that are unevenly distributed across space and time. The approach can reveal important events and location-specific trends in weather, traffic, and environmental data, such as road segments where congestion increases significantly during rainfall.
Event data warehouse

Event Data Warehouse

Spatial high-utility frequent pattern mining

Spatial High-Utility Frequent Pattern Mining

Distributed and Cyclically Evolving AI Infrastructure

AI models, data, and computing resources distributed across cloud, edge, and organizational environments are securely and efficiently coordinated. Through federated and continual learning, data and evaluation results obtained in real-world environments are fed back into model updates, enabling AI systems to improve through repeated cycles of use, evaluation, learning, and deployment.

Research Topics

  • Federated Learning for Heterogeneous Edge Environments
    AI models are collaboratively trained across multiple edge devices without directly sharing private local data. The approach enables vehicles, smartphones, and IoT devices with different computing capabilities and network conditions to jointly improve models while supporting continuous local adaptation.
  • Cyclically Evolving AI Model Framework
    Data, learning outcomes, and evaluation results are circulated among model providers, service providers, and users to continuously improve AI models. One application is a cyclically evolving AI dashboard-camera system that connects in-vehicle edge devices, office edge servers, and the cloud to update driving-risk prediction models according to actual usage conditions.
  • AI Frameworks and Smart Service Applications
    An integrated AI framework supports data collection, model development, deployment, evaluation, and continuous updating. By combining MLOps and distributed learning infrastructure, the framework can support the development and operation of smart mobility, driver assistance, and community safety services.
    Reference: National Institute of Information and Communications Technology, xData Platform and DCCS Testbed
Federated learning for heterogeneous edge environments

>Federated Split Learning for Heterogeneous Edge Environments.

Continually evolving AI dashboard-camera system

Cyclically Evolving AI Dashboard-Camera System

On-the-Air Federated Learning, which integrates wireless communications with federated learning.