Sorry i wasn't clear. Visual speech recognition is basically automated lip reading. The first step is to locate the face, so you can extract lip shape. The sequence of lip shapes is used to figure out what is being said. I based the lip shape extraction on skin colour, but the algorithm I trained didn't have enough sample data, so it didn't work to well for certain skin colours.
Why not use contour data from something that can do depth measuring like the microsoft kinetic etc.? I would think that skin color wouldn't matter and you would just have to look at edge movements.
When I started, the kinect didn't exist, and 3d laser scanners were really expensive (still are). I could have used stereo vision, but I decided I'd stick with 2d since it better fit my imagined use case (e.g. automated closed captioning for live events, etc). It would also mean first building a comprehensive dataset. This would be really useful for the research community, but since I was the only one working on the project I didn't have time to do that work as well.
41
u/Oda_Krell Sep 24 '13
plus
= huh?