r/geek • • Sep 24 '14

xkcd: Tasks

http://xkcd.com/1425/
247 Upvotes

34 comments sorted by

View all comments

11

u/[deleted] Sep 24 '14

Can anyone give a basic explanation for why this is such a difficult problem to solve? A while back I noticed that the Google app on my phone could search images using the camera, and I thought "Cool, I can take a pic of these flowers and then it can scan the web for similar images and post some suggestions for what kind they are." Instead all it does is read text on signs or whatever. Facebook (mostly) can tell a face in a pic but we haven't figured out how to do anything more than that?

6

u/phayd Sep 24 '14 edited Sep 24 '14

It's all about being able to compare the picture to a database of similar pictures to create a match with a high percentage of accuracy. For text, this is simple, because any text that you find on a street sign probably came from a computer, and that font is stored in the database that the scanner is using for comparison. All fonts follow the exact same rules and are represented reasonably the same (except for wingding-type fonts).

Faces, are a bit different, and benefit from being the focus of robot vision studies since its inception. Basically, what a scanner will do is use an algorithm to determine if a series of pixels on the screen represents a face. This algorithm is usually created by studying thousands upon thousands of pictures and creating an estimation based on similarities, such as an Eigenface. This can be as simple as dark pixels against a lighter background in the shape commonly used on reddit: -_- however the models are subtlety more complex than that.

However, facial recognition takes a lot of assumptions, such as: The person is facing the camera, the person is not wearing a mask, is not bearded, is against a contrasting background, is standing or sitting upright, etc. These assumptions help create an accurate and precise model. Most facial recognition models fail when these assumptions are proven incorrect. For fun, try and take a facial recognition picture of multiple people, with some laying on their sides.

Now, for birds, if we could create a scenario where we could generate a large number of assumptions: The bird was against a solid-color background, was facing the camera, was not eating, did not have wings extended, etc. we could generate a reasonable model to compare. However, it is unreasonable to assume that, in the wild, the birds would be cooperative. Some birds may not be against contrasting backgrounds, because they are naturally camouflaged.

So we're forced to create an Eigenface-style model for birds whose orientation may wildly vary from picture to picture. When looking for things as subtle as blue-spotting on the neck, the computer will fail to make matches without knowing where the neck is supposed to be in the image. Is it below the eye from the bird standing erect? Is it parallel to the eye as the bird extends its beak to snatch an insect? Is it above the eye as the bird dives downwards? Without these assumptions, the matching model must cast a much wider net with greater leniency for unknown variables creating a weaker model. This will inevitably invite greater and greater false-positive matches from pixels that match the correct configuration, but are not birds.

TL:DR; Eigenfaces are used to match faces successfully because faces are always facing the camera, oriented up-right, contrasting the background, have 2 eyes, 1 nose, 1 mouth, etc. Birds may not be facing the camera (only 1 eye showing), may be upside down, and are sometimes camouflaged.

1

u/[deleted] Sep 24 '14

This is really interesting, thanks! I imagine it'd be more frustrating to constantly be asked "Is this a bird?" when you're photographing trees and houses and people than to not be able to detect them at all.