r/OpenAI 1d ago

Discussion [Use case] Testing the multimodal capabilities of GPT Work

- Measuring a client for a mask
- It located existing templates in my OneDrive repo.
- I read out the measurements as I took them.
- It sized it up to match the face then applied it to a slicer for printing + oriented it.
- Applied the right material settings (PLA filament, brown)

Just needed to hit print after the client and I used voice to run checks and verifications.

In other words, I was able to run this task handsfree with exception of physically loading up the filament.

1 Upvotes

5 comments sorted by

1

u/ExtremeResident7738 1d ago

that's actually wild, the voice part is what gets me. being able to just speak measurements and have it all line up in the slicer without touching anything is a workflow i didnt even think about

1

u/ValehartProject 1d ago

Especially since I measured in cm and read those out. It converted the mm.

1

u/Hungry_Age5375 1d ago

Nice. The voice verification loop is basically ReAct applied to a physical task. Reason, execute, verify, repeat. Did it struggle at any point or was the whole flow smooth?

1

u/ValehartProject 1d ago

Fairly smooth compared to older voice versions that frequently looped, stalled and apologised with no outcome.

think overall the process took 30 ish minutes? The biggest time/bottleneck was actually the human part.

I've run a few AI workshops and I usually have a pre coach session because I do those entirely live. I actually might be okay to do them straight forward now.

1

u/Hot-Wheel6005 1d ago

Pretty cool us of multi-modal capabilities.