Researcher · Taiwan

Chieh-Yu Pan

Multimodal AI and Computer Vision

I study how visual information can be combined with language and structured domain knowledge for fine-grained prediction. My work spans multimodal learning, RGB-D vision, visual grounding, and structured AI systems, with applications in nutrition analysis, social robotics, and applied computer vision.

Academic background

2023–2026

National Cheng Kung University

M.S. in Computer Science and Information Engineering

GPA 4.06/4.30 · Master’s thesis: KARINA · Advisor: Prof. Wei-Ta Chu

Tainan, Taiwan

2025–2026

KTH Royal Institute of Technology

Exchange Studies in Computer Science, 37.5 ECTS

Image Analysis and Computer Vision · Big Data in Media Technology · Social Robotics

Stockholm, Sweden

2019–2023

National Cheng Kung University

B.S. in Engineering Science

GPA 3.95/4.30

Tainan, Taiwan

Areas of interest

My research centers on multimodal systems that connect visual evidence with language, knowledge, and structured reasoning. I am particularly interested in methods whose intermediate representations and failures can be examined rather than treated as opaque outputs.

01

Knowledge-grounded multimodal learning

Combining visual features with language and domain knowledge so that models can reason about fine-grained, semantically meaningful evidence.

Vision-language models · RGB-D fusion · Multimodal alignment
02

Fine-grained visual understanding

Localizing and interpreting small or compositional visual regions, particularly when direct supervision is limited or costly.

Segmentation · Visual grounding · Object detection
03

Structured AI systems

Integrating multimodal perception with planning, memory, retrieval, and bounded tool use for systems that operate under real-world constraints.

LLM agents · Persistent memory · Tool use · RAG

KARINA

Knowledge-Augmented Representation for Ingredient-Level Nutrition Analysis from Food Images

Research question

Can ingredient-level knowledge and spatial grounding improve the estimation of calories, mass, fat, carbohydrates, and protein from food images?

Method

KARINA combines RGB-D visual features with ingredient descriptions and qualitative nutritional profiles generated by a large multimodal model. Mask-based Ingredient-level Augmentation uses automatically generated Grounded SAM masks to align each description with its corresponding food region without additional manual annotation.

My contribution

I conducted this work as my master’s thesis under the supervision of Prof. Wei-Ta Chu. As the graduate researcher and first author, I designed and implemented the framework, carried out evaluation across datasets and ablation studies, analyzed the results, and released the public implementation. The paper is authored by Chieh-Yu Pan and Wei-Ta Chu.

Applied systems work

Selected projects showing experience in multimodal perception, constrained reasoning, system integration, and empirical evaluation.

2025–Present

KTH Royal Institute of Technology

Proactive Multimodal Social Robot

Focus
A Misty II prototype that detects opportunities to offer support before an explicit user request.
Methods
Visual and speech perception, structured LLM planning, three-tier persistent memory, and a bounded executor with consent-aware motion control.
Evidence
Evaluated in ten scripted multimodal cases and against AutoMisty’s 28 published benchmark tasks.
2025

TSMC Intelligent Manufacturing Workshop

Emergency Dispatch after AMHS Failure

Focus
A constrained dispatch problem involving 40 WIP lots, 20 carts, variable capacities, and a 50-location transfer-time matrix.
Methods
Queue-time, capacity, and routing constraints were formalized into a dispatch heuristic within a five-day cross-disciplinary project.
Evidence
Completed both evaluated settings without Q-time violations; team placed 2nd.
2022

DIGI+ Talent Program

Invoice Recognition and ERP Automation

Focus
A computer-vision pipeline for extracting invoice fields at an operational scale of approximately 55,000 invoices per month.
Methods
YOLOv5, OCR, iterative crop refinement, and downstream RPA/ERP integration.
Evidence
Achieved 96% invoice-level exact-match accuracy, with an estimated 78% reduction in processing time.

Research record

  1. 2026

    KARINA: Knowledge-Augmented Representation for Ingredient-Level Nutrition Analysis from Food Images

    Chieh-Yu Pan and Wei-Ta Chu

    Proceedings of the 2026 IEEE Conference on Artificial Intelligence (CAI), pp. 1–6

    First author
  2. 2022

    Integrating Virtual Reality Technology and Simulation-Based Learning to Enhance Student Learning Performance and Engagement

    Chieh-Yu Pan, Xiang-Ning Lu, Yu-Ping Cheng, Xiang-Yu Li, and Yueh-Min Huang

    Proceedings of the 2022 Taiwan Academic Network Conference (TANET), pp. 1268–1273

    In Chinese · Oral presentation
  3. 2022

    Enhancing Learning Performance of Engineering Students in Virtual Reality Environment

    Chieh-Yu Pan, Yu-Ping Cheng, and Yueh-Min Huang

    20th International Symposium on Novel and Sustainable Technology (ISNST), Tainan, Taiwan

    Short paper · Excellent Poster Presentation Award

Additional experience

Teaching and research service

2023–2025
Teaching Assistant, Introduction to Artificial Intelligence

Designed assignments and scalable grading workflows, held office hours, and supported more than 300 students.

2023–2025
Research Computing Administrator

Administered seven Linux servers and 22 NVIDIA GPUs for 15 researchers, including system upgrades and training-job troubleshooting.

Selected recognition

  • 2025–2026Transnational Study and Research Scholarship Grant
  • 20252nd Place, TSMC Intelligent Manufacturing Workshop
  • 2023NSTC Research Creativity Award
  • 2022Excellent Poster Presentation Award, 20th ISNST
  • 2022Outstanding Award, DIGI+ Talent Program

Technical methods

Programming & ML
Python, C/C++, PyTorch, Hugging Face Transformers, OpenCV
Computer vision
VLMs, RGB-D fusion, segmentation, detection, multimodal alignment
Systems & agents
Linux, NVIDIA GPU administration, structured LLM planning, persistent memory

Get in touch

For research collaboration, technical discussion, or professional inquiries, please feel free to contact me.