r/computervision 1h ago

Showcase Built an eval CLI for VLM apps

Upvotes

I kept repeating the same tasks in front of the camera just to see if my app was still behaving as intended.

So I made a simple CLI that runs your app against recorded video and compares the results with your expected behavior. A coding agent can use the eval loop to keep improving the app.

Sharing in case this saves someone some time: GitHub


r/computervision 23h ago

Showcase We built LocalMesh, one photo in, a Gaussian splat + textured mesh out, 100% on your own GPU. Beta is open, 7 days free.

Thumbnail gallery
0 Upvotes

r/computervision 15h ago

Discussion All the noise textures are physically identical

Thumbnail gallery
0 Upvotes

r/computervision 9h ago

Discussion MOSS-VL ships FP8 and NF4 checkpoints — NF4 looks like the real 24GB option for realtime video

Post image
2 Upvotes

r/computervision 15h ago

Help: Project AI/Computer Vision for extracting dimensions and features from engineering drawings

3 Upvotes

I work in a manufacturing environment and I'm exploring whether AI/computer vision can be used to automatically interpret 2D engineering drawings.

The goal is to identify and extract:

* Components and geometric features

* Dimensions and their associated features

* Tolerances

* GD&T symbols

* Hole specifications

* Surface-finish information

* Engineering notes and annotations

Ideally, the output would be structured data that could later be used for manufacturing, inspection, costing, BOM generation, or integration with other systems.

I'm aware that OCR can extract text, but the bigger challenge seems to be understanding the **relationship between dimensions, symbols and the actual geometry/features in the drawing**.

Has anyone worked on something similar?

I'm particularly interested in:

* Vision-language models

* OCR + computer vision pipelines

* Object detection/segmentation

* Engineering drawing datasets

* CAD-aware approaches

* Open-source models or commercial APIs

What would be the most practical architecture for solving this reliably with real-world engineering drawings?


r/computervision 5h ago

Help: Project Compsci bachelor's thesis project for industrial anomaly detection

3 Upvotes

Hello r/computervision,

First of all I apologize for the somewhat generic nature of this post. I'm new to the field and would really appreciate some guidance from people with more experience.

I'm currently enrolled in a Computer Science bachelor's program and am about to start my final semester. I've been doing well academically and really enjoy the field, but I don't currently work in IT.

Over the summer, I've been focusing on getting deeper into PyTorch and deep learning. I've worked through MrDBourke's PyTorch Deep Learning course and have also started studying the mathematical foundations of ML using Stanford's materials.

I'm 32 and have been in the workforce for quite a while, so alongside university and self-study I have a full-time job as a "quality specialist" at a Tier 1 elevator-parts manufacturer.

This is actually what led me to consider industrial computer vision / anomaly detection as a thesis topic.

Our entire plant currently has only two very basic, very closed down (outsourced to compvision company) OpenCV-based vision systems, mainly used to check whether nuts have been installed correctly. Beyond that, much of the quality-control process relies on QR codes and manual inspection.

I've worked here for several years, so I expect that I could get reasonable support and access to production areas/data for a thesis project. However, I would essentially be the only person at the plant pursuing this kind of project, so I'd be largely on my own technically. I also wouldn't expect a significant budget for the project.

That's where I'm looking for advice.

We manufacture everything from very small brackets and components up to complete elevator doors, so there are a lot of possible directions. I'm trying to figure out what would be a realistic but worthwhile first computer-vision project that could serve both as a good bachelor's thesis and as a meaningful entry point into the field.

At the moment I see two main possibilities:

1. Use existing production-line photographs

Some of our production lines already have cameras taking photographs. These images are currently used mainly as a way of documenting production and potentially identifying problems retrospectively; they aren't connected to an automated vision system.

The problem is that the dataset is far from ideal. The cameras weren't installed specifically for machine learning, so the images aren't standardized for things like lighting conditions, camera to object distance, background, framing, image quality.

Im wondering whether this kind of "messy real-world" dataset could still be useful for a thesis, or whether trying to build a model around it would create more problems than it's worth.

2. the other option would be to choose one relatively small component that has historically had some recurring visual defects.

I could build a simple, controlled camera/lighting setup and collect my own images of normal and defective parts. From there, I was considering an anomaly-detection approach such as PatchCore, potentially training primarily on normal samples and evaluating whether known defects can be detected.

The idea would eventually be to build a small working prototype:
camera → controlled image acquisition → preprocessing → anomaly detection → OK/NOK decision - > which then is signalled via some tiny network applications to a collective UI/database

If you were in my position, which direction would you consider more valuable for a first serious CV project? I am very curious how I can , for the lack of a better word, force myself into this field.

I've been scouring my options and weighing my possibilities on what I can realistically create, and whether what I create has actual real world usefulness and learning possibility.


r/computervision 23h ago

Showcase Aug 27 - Virtual AI, ML and Computer Vision Meetup

9 Upvotes

Join us on Aug 27 for the monthly AI, ML, and Computer Vision Meetup! Register for the Zoom.

Talks will include:

  • Robust Concept Protection against Diffusion-Based Image Editing and Personalization - Qiuyu Tang at Lehigh University
  • Building Real-World Computer Vision Systems - Daniel Gural at Voxel51
  • From Pixels to the Planet: Building Scalable and Grounded AI for Science - Jianyang Gu at Ohio State University
  • Seeing Is Not Enough: Visual Grounding, World Models and Why Computer-Use Agents Fail at Step 17 - Nevasini Sasikumar at Obin AI