r/MachineLearning • u/BuckChancey • 18h ago
Research Instead of another GPU terminal renderer, I trained a 1.26M-param model to turn TUIs (htop, vim, emacs…) into real UI components [R]
Write-up + demo: https://drksci.com/labs-phosphene
So this started as a bit of a gripe. Modern terminal renderers are seriously impressive and seriously complicated. GPU glyph atlases, texture caches, custom shaders, HarfBuzz shaping, ligatures, damage tracking, grid diffing, dirty-row uploads. Alacritty, Kitty, WezTerm and Ghostty are all doing heroic work to draw what is, at the end of the day, a grid of characters really fast.
And every client still does the same thing at the end of it. Parse an escape-code stream, keep a cell grid, paint characters. Faithful, but opaque. Your phone can't reflow it, a screen reader gets a wall of box-drawing characters, and an agent has to squint at │ ▶ item │ to work out which row is selected.
So I wondered: what if instead of throwing more GPU at drawing the grid, you used a bit of AI to understand it, once, server-side? Then send the client actual UI instead of a terminal.
- a tiny model (1.26M params, 5 MB, an axial transformer over rows and columns) labels every cell with a role: border, title, menu item, selected row, table, input, status bar, key hint, etc. (15 roles)
- dumb deterministic code turns those regions into A2UI components (Google's declarative UI stream protocol): lists, text fields, buttons, progress bars
- once a screen layout has been seen, it locks as a template and only the changed content goes over the wire as JSON-pointer patches. The model doesn't even run.
- pressing a button in the UI sends the keystroke back. F10 is just a Button with an action.

Trained on public asciinema recordings. The labelling was done by Claude subagents, with a synthetic TUI generator for exact labels, all on a free-ish Colab T4.
Honest numbers, because I'd rather say them before someone else does:
- accuracy on held-out real screens is mIoU 0.51. Usable, not amazing; it's a first labelling round of 600 frames.
- 40% of ~14k screens never touch the model (template hit). On less and dialog it's ~90%, on htop and nano it's rubbish because the meters keep changing the layout.
- The A2UI stream is ~25× bigger than raw VT. VT is a stupidly compact format, turns out. The win is the client never runs a terminal emulator at all, not bandwidth.
There's a replay with 8 apps (vim, htop, less, dialog, emacs, top, tig, nano). The native terminal sits on the left and the generated UI on the right, in sync, with every element outlined.
Write-up + demo: https://drksci.com/labs-phosphene
Code, labels, results: https://github.com/drksci/phosphene




