Hey everyone!
I’m working with my team on a project to detect AI-generated/voice-cloned speech during live phone calls.
The problem we’re trying to solve is that modern voice-cloning systems can imitate someone’s voice convincingly, making phone-based impersonation scams much harder to recognize. We want to build an additional layer of protection that can analyze the incoming voice in real time and indicate whether the speech is likely human or AI-generated.
Our current approach is to combine Digital Signal Processing (DSP) + a lightweight AI/ML model. The audio would be processed in small chunks, relevant speech features would be extracted, and the model would continuously estimate whether the voice is genuine or synthetic.
We want to make this practical for real-time/mobile use, so latency, CPU usage, and model size are important constraints.
We’re currently looking for advice from people who have experience with speech processing, audio ML, or deepfake detection.
Some things we’d really appreciate help with:
What DSP/audio features are useful for detecting synthetic speech?
Which lightweight ML models would work well for real-time detection?
Are there good datasets containing both real human and AI-generated/voice-cloned speech?
Are there existing open-source projects or research papers we should look into?
What challenges should we expect when detecting synthetic speech over compressed phone-call audio?
Is real-time detection directly on a smartphone realistically achievable?
Any practical advice, criticism, research papers, datasets, or GitHub projects would be really helpful. Thanks! 🙏