← Projects

Group coursework · University of Bristol · 2026

Causal Audio-Visual Speech Enhancement with Mamba

The project built a real-time audio-visual speech enhancement system on a causal Mamba backbone and tested it under noisy and disrupted visual input. I co-led the six-person team. The overall system design, the integration of all components, and the training and evaluation were my responsibility. The report’s work allocation also lists my parts of the model: the visual confidence estimator, the Scout lookahead design, the AudioNeed amplifier and its later removal, the multi-task loss and the multiplicative fusion cascade. In addition, I built the three-channel cocktail-party mixing protocol and ran the bug audit, the configuration ablations and the per-SNR contribution experiment. The report was not peer reviewed, and its results belong to all six authors.

View code on GitHub ↗

Course report

Trust is All You Need: Real-Time Audio-Visual Speech Enhancement with Reliability-Aware Causal Mamba

Billy Hou, Dominic Jackson, Shunjie Qiu, Ronny Tu, Fan Xu, Zhenning Zhao

Supervised by Dr F. Karameh
School of Electrical, Electronic, and Mechanical Engineering · University of Bristol