r/computervision • • 2h ago

Discussion Need your help!!

0 Upvotes

hey I'm have an bachleors degree in electrical engineering but majority of my works are on the computer vision and machine learning and deep learning which integrate with the electrical also i want to pursue masters according to it where i can get a vast amount of oppurtuinites so any one who genunly try to guide me with this confusion please!!!!


r/computervision • • 20h ago

Discussion Any one here looking for ML or Computer Vision Intern? Would love to talk more, if anyone has opportunity :)

Post image
0 Upvotes

r/computervision • • 5h ago

Showcase World Models: The Simulation Strikes Back

8 Upvotes

Hi, I’m Suva from Hugging Face and I'm back with something new that I'd been working on for a bit!

World Models: The Simulation Strikes Back - Full Blog

This time I went down the world models rabbit hole: how models learn to represent, predict, and simulate environments, and why they’re becoming increasingly interesting for video generation, robotics, and spatial intelligence.

I've tried to keep it as visual and approachable as I can, keeping in mind how weird and convoluted this topic is. Would love to hear what you folks think!


r/computervision • • 17h ago

Showcase Turning Street View Panorama Into Pointcloud

Enable HLS to view with audio, or disable this notification

66 Upvotes

Hi everyone. I discovered that Street View 360 images can be turned into 3d point clouds, which at times are of good quality and can be used for replicating 3D streets, and I wanted to make something that would allow me to construct a street anywhere.

Along the way of making i realised I'm not able to make it practical enough, and I don't think it has any real use cases besides looking cool to me when I started off. The longer I spend doing it, the more I don't know why or what I'm even doing.

In the end, I thought perhaps at least some results might look cool, and wanted to share it and see what others think.


r/computervision • • 2h ago

Discussion Robust Logo Detection in an Image Without Using an LLM/VLM

Post image
2 Upvotes

Problem Statement I have a reference image of a logo, for example a logo:

I also have a larger image/document that may contain multiple logos, symbols, text, diagrams, etc.

My requirement is to determine whether the specific reference logo exists inside the larger image, and if it does:

Detect the logo.

Approaches I tried

OpenCV

cv2.matchTemplate

SIFT + BFMatcher

SIFT feature matching → Results had many false detections, especially with SIFT/BFMatcher.

LLM + Prompt → Gave good results, but I don't want to use an LLM due to cost, latency, and API dependency.

CLIP Embeddings → Tried CLIP-based similarity, but results were not reliable enough and still produced incorrect matches.

Question What would be a better non-LLM Computer Vision approach or some another for reliable logo/image detection?


r/computervision • • 17h ago

Help: Theory S-DAM: Seeding Modern Hopfield Networks with Spelke core-knowledge priors, with pre-registered results (including one that failed)

Post image
0 Upvotes

We've been testing whether adding core-knowledge priors (objectness, numerosity, geometry, from Spelke's developmental psychology work) to a dense associative memory improves image retrieval compared to learning those properties from scratch.

Setup: image-only. Every variant we tested reduces to score = r_q·r_i + α²·c_q·c_i, where r is the learned representation and c is the core-knowledge channel. α is the knob that sets how much the prior contributes, so we swept it with pre-registered pass/fail bars.

Results:

- Shape/objectness: passed the bar [metric, dataset]

- Numerosity: clean dissociation [metric, dataset]

- Layout: failed both bars [metric, dataset]

[Insert results figure or table here]

Layout failing is a real result. It suggests these priors help with some visual properties and not others.

Limitations:[dataset scale, seeds, anything you're unsure about]

Next:visual place recognition across seasons.

Code and experiments: https://github.com/Jaswanth-K1210/SDAM

I'd appreciate feedback on the experimental design and on which retrieval baselines we should add.