Yichi Zhang

UMich_Yichi.jpg

Hello! My name is Yichi Zhang, and I am doing multimodal research at ByteDance Seed. My research focuses on real-time multimodal agentic systems that integrate perception, reasoning, and interaction across streaming vision and audio. I aim to build proactive, full-duplex agents that can continuously understand and act in both digital and physical worlds, collaborating with people naturally and at the right moment.

During my Ph.D., I also worked on embodied AI agents that follow natural-language instructions and collaborate with people in situated, interactive environments. I received my Ph.D. in Computer Science and Engineering from the University of Michigan, where I was advised by Professor Joyce Chai as a member of the SLED lab. In 2023, I led Team SEAGULL to win first place in the inaugural Amazon Alexa Prize SimBot Challenge.

Before joining UMich, I obtained my Master’s in Information and Communication Engineering at Tsinghua University in 2020, advised by Professor Zhijian Ou. In 2019, I worked with Professor Zhou Yu as a visiting scholar on task-oriented dialog systems. I received my Bachelor’s in Electronic Information Science and Technology from Tsinghua University in 2017.

News

Aug 05, 2026 Excited to release SeedRealtime, a native audio-visual full-duplex LLM. As a core contributor, I worked on visual-based proactive interaction and duplex RL training, enabling the model to continuously watch, listen, and respond at the right moment.
Jul 01, 2026 We released the Seed2.0 Model Card, presenting a frontier model for real-world complexity. I contributed to its streaming-video and omni-modal understanding capabilities.
Apr 15, 2026 We released the Seedance 2.0 paper, presenting a unified multimodal audio-video generation model. I contributed to its omni-modal data pipeline.
Mar 21, 2026 We released the Seed1.8 Model Card, introducing a generalized agentic model for real-world scenarios. I contributed to its streaming-video understanding and proactive visual interaction capabilities.
May 23, 2025 Excited to share that I’ve joined the Multimodal Interaction & World Model team at ByteDance! Looking forward to tackling new challenges ahead. I’m open to collaborations and mentoring interns — feel free to reach out!

Publications

2025

  1. arXiv
    proassist.png
    Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
    Yichi Zhang, Xin Luna Dong , Zhaojiang Lin , Andrea Madotto , Anuj Kumar , Babak Damavandi , Joyce Chai , and Seungwhan Moon
    In arXiv , Jun 2025
  2. COLM
    colm25_basis.png
    Bootstrapping Visual Assistant Development with Situated Interaction Simulation
    Yichi Zhang, Run Peng , Lingyun Wu , Yinpei Dai , Xuweiyi Chen , Qiaozi Gao , and Joyce Chai
    In Second Conference on Language Modeling , Oct 2025

2024

  1. CVPR
    cvpr24_groundhog.png
    GROUNDHOG: Grounding Large Language Models to Holistic Segmentation
    Yichi Zhang, Ziqiao Ma , Xiaofeng Gao , Suhaila Shakiah , Qiaozi Gao , and Joyce Chai
    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , Jun 2024
  2. COLM
    colm24.png
    Autonomous evaluation and refinement of digital agents
    Jiayi Pan , Yichi Zhang, Nicholas Tomlin , Yifei Zhou , Sergey Levine , and Alane Suhr
    In First Conference on Language Modeling , Oct 2024

2023

  1. EMNLP
    emnlp23_illusion.png
    Grounding Visual Illusions in Language: Do Vision-Language Models Perceive Illusions Like Humans?
    Yichi Zhang, Jiayi Pan , Yuchen Zhou , Rui Pan , and Joyce Chai
    In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , Dec 2023
  2. EMNLP Findings
    emnlp23_wtag.png
    Can Foundation Models Watch, Talk and Guide You Step by Step to Make a Cake?
    Yuwei Bao , Keunwoo Yu , Yichi Zhang, Shane Storks , Itamar Bar-Yossef , Alex Iglesia , Megan Su , Xiao Zheng , and Joyce Chai
    In Findings of the Association for Computational Linguistics: EMNLP 2023 , Dec 2023

2022

  1. EMNLP
    emnlp22_danli.png
    DANLI: Deliberative Agent for Following Natural Language Instructions
    Yichi Zhang, Jianing Yang , Jiayi Pan , Shane Storks , Nikhil Devraj , Ziqiao Ma , Keunwoo Yu , Yuwei Bao , and Joyce Chai
    In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , Dec 2022

2021

  1. ACL Findings
    acl21_hitut.png
    Hierarchical Task Learning from Language Instructions with Unified Transformers and Self-Monitoring
    Yichi Zhang, and Joyce Chai
    In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , Aug 2021
  2. EACL
    eacl21_ardm.png
    Alternating Recurrent Dialog Model with Large-scale Pre-trained Language Models
    Qingyang Wu , Yichi Zhang, Yu Li , and Zhou Yu
    In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume , Apr 2021
  3. EMNLP Findings
    emnlp21_trip.png
    Tiered Reasoning for Intuitive Physics: Toward Verifiable Commonsense Language Understanding
    Shane Storks , Qiaozi Gao , Yichi Zhang, and Joyce Chai
    In Findings of the Association for Computational Linguistics: EMNLP 2021 , Nov 2021
  4. Applied Sciences
    ecrf.png
    Elastic CRFs for Open-Ontology Slot Filling
    Yinpei Dai , Yichi Zhang, Hong Liu , Zhijian Ou , Yi Huang , and Junlan Feng
    Applied Sciences, Nov 2021

2020

  1. AAAI
    aaai20_fig.png
    Task-oriented dialog systems that consider multiple appropriate responses under the same context
    Yichi Zhang, Zhijian Ou , and Zhou Yu
    In Proceedings of the AAAI Conference on Artificial Intelligence , Jan 2020
  2. ACL
    acl20_parg.png
    Paraphrase Augmented Task-Oriented Dialog Generation
    Silin Gao , Yichi Zhang, Zhijian Ou , and Zhou Yu
    In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Jul 2020
  3. EMNLP
    emnlp20_labes.png
    A Probabilistic End-To-End Task-Oriented Dialog Model with Latent Belief States towards Semi-Supervised Learning
    Yichi Zhang, Zhijian Ou , Min Hu , and Junlan Feng
    In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , Nov 2020
  4. INTERSPEECH
    is2020_dasi.png
    Improved Learning of Word Embeddings with Word Definitions and Semantic Injection.
    Yichi Zhang, Yinpei Dai , Zhijian Ou , Huixin Wang , and Junlan Feng
    In INTERSPEECH , Nov 2020