FlashLabs · Shanghai

Zhenghua Bao

包正华

Multimodal AI Engineer · FlashLabs

I work in the research department at FlashLabs in Shanghai, where I co-lead Chroma — a real-time end-to-end speech model with personalized voice cloning. My interests sit at the intersection of speech, language, and reasoning, with ongoing work on retrieval-augmented LLMs (IndexRAG). I completed MSc degrees in CS (AI) and Internet & Web-based Systems at TU Darmstadt.

§ 01 / RESEARCH

I build systems that reason and speak. My current work centers on two threads: end-to-end speech models that can clone a speaker from seconds of audio (Chroma), and retrieval architectures that pre-compute cross-document inference at index time (IndexRAG). Earlier work explored multi-agent reinforcement learning with graph neural networks — specifically, generalization across unseen network topologies — and test-time adaptation for EEG foundation models.

speech generation voice cloning RAG multimodal AI multi-agent RL
§ 02 / PUBLICATIONS
2026

OrcaRouter: A Production-Oriented LLM Router with Hybrid Offline-Online Learning

Zhenghua Bao, Fengya Tian, Chris Zhang, Zhenjun Chen, Xile Ma, Yi Shi

Technical Report arXiv First Author

2026

IndexRAG: Bridging Facts for Cross-Document Reasoning at Index Time

Zhenghua Bao, Yi Shi

EMNLP 2026 Under Review First Author

2026

FlashLabs Chroma 1.0: A Real-Time End-to-End Spoken Dialogue Model with Personalized Voice Cloning

Tanyu Chen*, Tairan Chen*, Kai Shen*, Zhenghua Bao*, Zhihui Zhang, Man Yuan, Yi Shi

arXiv Preprint Co-First Author

2025

NeuroTTT: Bridging Pretrain–Downstream Task Misalignment in EEG Foundation Models via Test-Time Training

Suli Wang, Yangshen Deng, Zhenghua Bao, Xinyu Zhan, Yiqun Duan

NeurIPS 2026 Under Review

2024

Towards Generalizability of Multi-Agent Reinforcement Learning in Graphs with Recurrent Message Passing

Jannis Weil, Zhenghua Bao, Osama Abboud, Tobias Meuser

AAMAS 2024Oral

§ 03 / NEWS
2026
OrcaRouter: A Production-Oriented LLM Router with Hybrid Offline-Online Learning
2026.05 Technical report released on arXiv.
2026.06 Open-source code reached 650+ GitHub stars.
2026.03 IndexRAG preprint released on arXiv.
2026
FlashLabs Chroma 1.0: A Real-Time End-to-End Spoken Dialogue Model with Personalized Voice Cloning
2026.01 Open-sourced and released as a preprint.
2026.03 Reached 540+ GitHub stars and 15k+ HuggingFace downloads.
2025.11 Joined FlashLabs in Shanghai as a Multimodal AI Engineer.
2025.11 Completed MSc in Computer Science (AI) at TU Darmstadt.
2025.09 NeuroTTT preprint released on arXiv.
2024.10 Completed MSc in Internet and Web-based Systems at TU Darmstadt.
2024
Towards Generalizability of Multi-Agent Reinforcement Learning in Graphs with Recurrent Message Passing
2024.02 Preprint released on arXiv.
2024.05 Accepted to AAMAS 2024.
§ 04 / CONTACT

Drop me a line — I'm always happy to chat about speech, retrieval, or anything between research and shipping.