Zhenghua Bao
Multimodal AI Engineer · FlashLabs
I work in the research department at FlashLabs in Shanghai, where I co-lead Chroma — a real-time end-to-end speech model with personalized voice cloning. My interests sit at the intersection of speech, language, and reasoning, with ongoing work on retrieval-augmented LLMs (IndexRAG). I completed MSc degrees in CS (AI) and Internet & Web-based Systems at TU Darmstadt.
I build systems that reason and speak. My current work centers on two threads: end-to-end speech models that can clone a speaker from seconds of audio (Chroma), and retrieval architectures that pre-compute cross-document inference at index time (IndexRAG). Earlier work explored multi-agent reinforcement learning with graph neural networks — specifically, generalization across unseen network topologies — and test-time adaptation for EEG foundation models.
IndexRAG: Bridging Facts for Cross-Document Reasoning at Index Time
EMNLP 2026 Under Review First Author
FlashLabs Chroma 1.0: A Real-Time End-to-End Spoken Dialogue Model with Personalized Voice Cloning
arXiv Preprint Co-First Author
Drop me a line — I'm always happy to chat about speech, retrieval, or anything between research and shipping.