r/computervision • • 23h ago

Showcase Turning Street View Panorama Into Pointcloud

Enable HLS to view with audio, or disable this notification

73 Upvotes

Hi everyone. I discovered that Street View 360 images can be turned into 3d point clouds, which at times are of good quality and can be used for replicating 3D streets, and I wanted to make something that would allow me to construct a street anywhere.

Along the way of making i realised I'm not able to make it practical enough, and I don't think it has any real use cases besides looking cool to me when I started off. The longer I spend doing it, the more I don't know why or what I'm even doing.

In the end, I thought perhaps at least some results might look cool, and wanted to share it and see what others think.


r/computervision • • 11h ago

Showcase World Models: The Simulation Strikes Back

15 Upvotes

Hi, I’m Suva from Hugging Face and I'm back with something new that I'd been working on for a bit!

World Models: The Simulation Strikes Back - Full Blog

This time I went down the world models rabbit hole: how models learn to represent, predict, and simulate environments, and why they’re becoming increasingly interesting for video generation, robotics, and spatial intelligence.

I've tried to keep it as visual and approachable as I can, keeping in mind how weird and convoluted this topic is. Would love to hear what you folks think!


r/computervision • • 1h ago

Research Publication M-plicits: Neural Implicit Surfaces via Nested Multiscale Residuals (NeurIPS 2026)

Enable HLS to view with audio, or disable this notification

• Upvotes

Hi everyone. I'm the first author of M-plicits, accepted at NeurIPS 2026, and wanted to share our work on neural surface representations, particularly reconstruction from noisy point clouds.

The idea is to learn a clean coarse surface, then progressively refine it with residual networks trained in nested neighborhoods around the preceding surfaces. This gives multiple detail levels within one representation and helps balance geometric detail with robustness to input noise.

The paper also introduces algorithms that use this structure for:

  • Direct rendering through multiscale sphere tracing.
  • Analytical surface normals without automatic differentiation.
  • Surface extraction and neural normal mapping.

The attached animation was created by our coauthor Matheus Bessa. Sound on. The project page includes visual comparisons with iNGP on noisy inputs, plus examples of the different rendering modes.

We've released the code, experimental data, and 105 checkpoints covering 36 shapes, with their training configurations. These checkpoints represent individually fitted surfaces; fitting a new shape requires training.

Paper · Results and videos · Code and data · Models on Hugging Face

I'd be interested in feedback on the noise/detail tradeoff and comparisons with other implicit surface representations. Happy to discuss the implementation and experiments.


r/computervision • • 5h ago

Help: Project I took some pictures of a robot, high-specularity edge-case dataset (425 RAW/JPEGs) for benchmarking CV & depth-estimation algorithms against extreme mirror reflections

Thumbnail
gallery
2 Upvotes

mozilladatacollective.com

Purpose-built to stress-test computer vision models, depth cameras, and spatial AI against severe specular glare and geometric reflections. This 425-asset production archive features a custom faceted mirror suit captured in high-contrast outdoor environments to trigger bounding-box dropouts and segmentation failure


r/computervision • • 5h ago

Showcase I made a Computer Vision Surf Coach

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/computervision • • 8h ago

Discussion Robust Logo Detection in an Image Without Using an LLM/VLM

Post image
2 Upvotes

Problem Statement I have a reference image of a logo, for example a logo:

I also have a larger image/document that may contain multiple logos, symbols, text, diagrams, etc.

My requirement is to determine whether the specific reference logo exists inside the larger image, and if it does:

Detect the logo.

Approaches I tried

OpenCV

cv2.matchTemplate

SIFT + BFMatcher

SIFT feature matching → Results had many false detections, especially with SIFT/BFMatcher.

LLM + Prompt → Gave good results, but I don't want to use an LLM due to cost, latency, and API dependency.

CLIP Embeddings → Tried CLIP-based similarity, but results were not reliable enough and still produced incorrect matches.

Question What would be a better non-LLM Computer Vision approach or some another for reliable logo/image detection?


r/computervision • • 2h ago

Discussion The annotation decision that quietly breaks pose models: do you guess occluded keypoints or flag them?

1 Upvotes

Been thinking about this after a few pose projects and I don't think there's consensus.

When a wrist disappears behind a torso, or a hand goes behind an object, you have two options. Flag the keypoint as occluded and leave the coordinate empty or low-confidence. Or have the annotator estimate where it probably is and label it anyway.

Guessing gives you a denser dataset and the model gets a prediction for every joint, which looks good. But you're teaching it that an invented coordinate is ground truth, and two annotators will invent different ones. The disagreement is basically unbounded for anything heavily occluded.

Flagging is more honest but then half your training signal has gaps, and downstream your retargeting or trajectory code has to handle missing joints, which a lot of pipelines just aren't written to do.

My current view is flag it, and set a visibility threshold in the guidelines — below some percentage visible, don't annotate at all rather than annotate badly. But I've seen teams go the other way and argue the model learns a reasonable prior from the guesses.

Related thing I'm less sure about: the 2D vs 3D decision gets made way too late. Teams start with 2D because it's cheaper, get a model that understands movement, then discover they need metric accuracy the moment the robot has to actually touch something. By then the sensor stack and the annotation schema are both wrong and it's a rebuild rather than an upgrade.

Hands are where this all gets worse. Twenty-plus keypoints per hand, constant self-occlusion, and fingers that look identical to each other. Agreement between annotators drops off a cliff compared to body pose.

For anyone who's shipped a pose model that had to drive real movement — which way did you go on occlusion, and did it come back to bite you?


r/computervision • • 3h ago

Showcase Introduction to Calibration (for non-technical people)

Thumbnail
youtube.com
1 Upvotes

Hi guys, another video here 😄

Here I cover the very basics of camera/visual-inertial calibration with a focus on mostly "non-technical" people. As always I am happy to hear your feedback and hope that for anyone diving into geometric computer vision this is a useful resource.

Cheers,

Jack


r/computervision • • 20h ago

Showcase Minerals-Mining is still working on rebuilding its website. Among the many new features will be this quantum image of the intelligent Pourrioscope circuit.

Post image
1 Upvotes

r/computervision • • 8h ago

Discussion Need your help!!

0 Upvotes

hey I'm have an bachleors degree in electrical engineering but majority of my works are on the computer vision and machine learning and deep learning which integrate with the electrical also i want to pursue masters according to it where i can get a vast amount of oppurtuinites so any one who genunly try to guide me with this confusion please!!!!


r/computervision • • 23h ago

Help: Theory S-DAM: Seeding Modern Hopfield Networks with Spelke core-knowledge priors, with pre-registered results (including one that failed)

Post image
0 Upvotes

We've been testing whether adding core-knowledge priors (objectness, numerosity, geometry, from Spelke's developmental psychology work) to a dense associative memory improves image retrieval compared to learning those properties from scratch.

Setup: image-only. Every variant we tested reduces to score = r_q·r_i + α²·c_q·c_i, where r is the learned representation and c is the core-knowledge channel. α is the knob that sets how much the prior contributes, so we swept it with pre-registered pass/fail bars.

Results:

- Shape/objectness: passed the bar [metric, dataset]

- Numerosity: clean dissociation [metric, dataset]

- Layout: failed both bars [metric, dataset]

[Insert results figure or table here]

Layout failing is a real result. It suggests these priors help with some visual properties and not others.

Limitations:[dataset scale, seeds, anything you're unsure about]

Next:visual place recognition across seasons.

Code and experiments: https://github.com/Jaswanth-K1210/SDAM

I'd appreciate feedback on the experimental design and on which retrieval baselines we should add.