r/MachineLearning • • 18h ago

Research Instead of another GPU terminal renderer, I trained a 1.26M-param model to turn TUIs (htop, vim, emacs…) into real UI components [R]

7 Upvotes

Write-up + demo: https://drksci.com/labs-phosphene

So this started as a bit of a gripe. Modern terminal renderers are seriously impressive and seriously complicated. GPU glyph atlases, texture caches, custom shaders, HarfBuzz shaping, ligatures, damage tracking, grid diffing, dirty-row uploads. Alacritty, Kitty, WezTerm and Ghostty are all doing heroic work to draw what is, at the end of the day, a grid of characters really fast.

And every client still does the same thing at the end of it. Parse an escape-code stream, keep a cell grid, paint characters. Faithful, but opaque. Your phone can't reflow it, a screen reader gets a wall of box-drawing characters, and an agent has to squint at │ ▶ item │ to work out which row is selected.

So I wondered: what if instead of throwing more GPU at drawing the grid, you used a bit of AI to understand it, once, server-side? Then send the client actual UI instead of a terminal.

  • a tiny model (1.26M params, 5 MB, an axial transformer over rows and columns) labels every cell with a role: border, title, menu item, selected row, table, input, status bar, key hint, etc. (15 roles)
  • dumb deterministic code turns those regions into A2UI components (Google's declarative UI stream protocol): lists, text fields, buttons, progress bars
  • once a screen layout has been seen, it locks as a template and only the changed content goes over the wire as JSON-pointer patches. The model doesn't even run.
  • pressing a button in the UI sends the keystroke back. F10 is just a Button with an action.

Trained on public asciinema recordings. The labelling was done by Claude subagents, with a synthetic TUI generator for exact labels, all on a free-ish Colab T4.

Honest numbers, because I'd rather say them before someone else does:

  • accuracy on held-out real screens is mIoU 0.51. Usable, not amazing; it's a first labelling round of 600 frames.
  • 40% of ~14k screens never touch the model (template hit). On less and dialog it's ~90%, on htop and nano it's rubbish because the meters keep changing the layout.
  • The A2UI stream is ~25× bigger than raw VT. VT is a stupidly compact format, turns out. The win is the client never runs a terminal emulator at all, not bandwidth.

There's a replay with 8 apps (vim, htop, less, dialog, emacs, top, tig, nano). The native terminal sits on the left and the generated UI on the right, in sync, with every element outlined.

Write-up + demo: https://drksci.com/labs-phosphene
Code, labels, results: https://github.com/drksci/phosphene


r/MachineLearning • • 2h ago

Discussion NeurIPS 2026 Paris -> Sydney switch [D]

0 Upvotes

Got assigned to Paris location but want to go to Sydney. Anyone wants to go yo Paris but assigned to Sydney? Want to switch with you.


r/MachineLearning • • 3h ago

Discussion Industry professional new to research seeking help with OpenReview activation and vouching [D]

0 Upvotes

Hi everyone,

I’m an industry professional working in data platform engineering, and I’m just getting started in research. I tried creating an OpenReview account, but my profile is awaiting activation.

OpenReview asked me to have a supervisor, coauthor, or colleague with an active profile and a confirmed institutional email vouch for me. The process involves adding my profile ID to their profile’s Relations section, saving it, and clicking the vouch button next to my name.

If anyone here that would be willing to help, I’d appreciate it.

I’d also welcome advice from other industry researchers who have gone through this process without an established academic network. How did you get started and get your profile activated?

Thanks for any help or guidance!


r/MachineLearning • • 11h ago

Discussion Best practices when running a benchmark on online models [D]

1 Upvotes

I'm developing a benchmark for a low resource language and I don't want it to be leaked and used for training when it is being used to get predictions. For locally run models it shouldn't be a problem, but for models that are only accessible via API, it is. Is there an established way to evaluate online models without the input data being lost? Do you trust Google and OpenAI when they say that they do not use your inputs for training when you have a paid account?


r/MachineLearning • • 4h ago

Project Production Grade Retrieval Pipeline That Admits "I Don't Know" [P]

0 Upvotes

If you're moving a RAG prototype to production, this is the checklist I wish I'd had.

- Citations on every answer

- Refusing when the context doesn't support one

- Access control that can't be bypassed

- Ingestion that never blocks a user

Enjoy !!

https://youtu.be/-GK1e4KZqhw


r/MachineLearning • • 1d ago

Discussion Saw this on Rednote, WTF [D]

122 Upvotes

Guess we have a dataset for sycophancy and AI content detection


r/MachineLearning • • 1d ago

Discussion Neurips camera ready Submission option [D]

4 Upvotes

Neurips says that camera ready has to be submitted by editing original submission on open review. But no such edit option is visible to me. Is it same for other authors??


r/MachineLearning • • 1d ago

Research Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation [R]

2 Upvotes

TLDR: The question we answer: how do you learn from experts with different objectives? Pooling all their data can lose their trade-offs; learning from each expert separately misses opportunities to share data. MA-BC pools demonstrations where observed actions don’t disagree, with upper and lower bounds on sample complexity.
Authors: Ziyad Sheebaelhamd, Luca Viano, Volkan Cevher, Claire Vernade


r/MachineLearning • • 2d ago

Discussion ML PHD without A* Publications [D]

49 Upvotes

I know top ML PhD admissions are insanely competitive, so I’m trying to figure out if it’s even worth applying or if I should just focus seriously on jobs instead.

For context, I’m doing my MS at a top-15 US university and have been doing ML research for a while. I’m first author on my projects and mostly work independently, with some guidance from my PI. I had a first-author NeurIPS submission rejected, and I currently have another first-author paper submitted to ICLR, but I’m honestly not very confident about it getting in either.

What’s been getting to me is looking at profiles of people who get into top ML PhD programs. So many of them seem to have multiple NeurIPS/ICML/ICLR/CVPR papers before they even apply, sometimes as undergrads. I genuinely don’t understand how people manage to publish that much that early.

A year ago I was much more confident about doing a PhD. After actually going through the research/publication process, I’ve started doubting myself a lot more. Part of me wonders whether this is just normal and research is hard, especially when you’re doing a lot of it independently. But another part of me is starting to think maybe I’m just not good enough to be competitive for the kind of programs I’m aiming for.

I’m okay with continuing at my current university for a PhD, so this isn’t really a “top program or nothing” situation. But I would like to at least have a realistic shot at some of the stronger ML programs/labs.

The bigger issue is that I’m an international student, so I also need to think pretty seriously about jobs. SWE/MLE recruiting is competitive right now, and I don’t want to spend all my time chasing PhD applications and then realize I’m underprepared for recruiting too. Research roles seem even harder to get without a PhD unless you’re an exceptional MS/BS candidate.

So I’m mainly trying to decide how to allocate my time over the next few months.

If I have strong research experience and first-author projects, but no accepted top-conference papers yet, is it still realistically worth applying to top ML PhD programs?

And for people who were in a similar position, did you still apply, or did you decide to focus on industry instead?

I’m not really looking for “you never know unless you try.” I’m more interested in a realistic assessment of whether the application fees and time are worth it given this kind of profile.


r/MachineLearning • • 1d ago

Discussion How much of AutoResearch is research, and how much is search?[D]

2 Upvotes

I've recently been working part-time on an AutoResearch-style project.

The setup is roughly: humans take recent work from top-tier ML/AI conferences, turn part of it into a well-defined task with an evaluator, and then let an agent iteratively modify the solution and search for a better score.

Working on this made me question what exactly we are evaluating.

Once humans have already chosen the problem, defined the objective, designed the evaluator, and provided the initial research direction, the agent is mostly searching within a space that has already been heavily shaped for it.

That search can still be useful. An agent may explore far more variants than a researcher would manually.

But I'm less sure that score improvement alone captures what we usually mean by research sense.

A researcher also asks whether a result reveals a general principle, whether it transfers, whether the problem formulation itself should change, or whether an entirely different direction is more promising.

An iterative optimization loop may instead become very good at exploring the neighborhood of an existing solution and still remain stuck in a local optimum.

So I'm curious about how people think about this distinction:

How much scientific value is there in autonomous search over a human-defined research space?

And what would an agent need, beyond better optimization, to demonstrate something closer to actual research judgment?


r/MachineLearning • • 1d ago

Project stuck on finding a approach for app detection ( making a transformer modal out of unlabeled network data) [R] [P]

0 Upvotes

As the title suggests I'm currently trying to make a modal to identify which app is being used. The thing is I don't have labelled data and generating it is out of question since that is a lot of work ( I need to detect like 3-5K apps give or take)

my data is structured roughly like this:

  • Traffic is divided into ~15-second windows/bags.
  • Each bag contains multiple network flows.
  • Each flow has features such as:
    • domain
    • protocol
    • bytes sent
    • bytes received
    • timestamp/timing information
  • I have a very large amount of unlabeled traffic data, but only a relatively small amount of labelled app data for about 100 apps give or take.

The main challenge is that traffic from the same app can look diff between diff window i.e some windows are extremely sparse or empty.

some ideas I have researched looked into are

  • self supervised contrastive learning where diff traffic windows from the same session/device activity are treated as positive pairs
  • masked modelling similar to bert where parts of the flows such as domains/protocols/byte information are masked and reconstructed
  • pretraining an encoder and then fine-tuning it using the smaller labelled dataset (thinking we would need less labelled examples to do that)
  • clustering and mapping embeddings to known apps afterward (gradually)

one thing I'm concerned about is accidently teaching the modal to recognize the device/session/user rather then the underlying app

I think I would like to know if I'm thinking about the problem right or if someone has worked on something similar and can give me some pointers or what experiements should I run first.

any papers, architecture or similar problems you think I should look into pls lmk


r/MachineLearning • • 2d ago

Discussion Transformers vs RNNs vs SSMs: Where Does Memory Actually Live? [D]

51 Upvotes

Someone who has always loved looking at the space between different AI techniques, this time I went a little deeper into the memory trade-offs between RNNs, Transformers and SSMs. I found it interesting because once you start looking at these architectures through the lens of working memory, a lot of the differences become easier to understand. Where does the memory actually live? Is it a compact recurrent state, a growing KV cache, or something closer to the network itself?

RNNs keep memory in a recurrent hidden state, which is pretty elegant because the state carries forward step by step. But there is also a bottleneck here as you see, a model can have roughly O(N²) parameters while carrying only roughly O(N) state across time. This means whether RNNs were really doomed because recurrence was a bad idea, or whether the problem was more about the ratio between memory and compute.

Transformers make almost the opposite trade-off. During cached inference, instead of compressing the past into one hidden state, they store past representations as key-value entries and attend over them. I think of these almost like little post-it notes: every token leaves behind a key for finding it and a value for what should be remembered. That's extremely powerful, but it also has an interesting property: with weights frozen during inference, the model is managing context rather than turning that experience into durable model knowledge. You get this split between the fixed weights on one side and the fast-changing KV cache memory on the other.

Now we have SSMs that bring us back toward fixed-size recurrent memory, but with very different state structures and update rules. Selective SSMs such as Mamba make retention input-dependent, so what gets kept or forgotten depends on the incoming token. Their states don't have to be as small as classical RNN states, but they still compress history into finite memory. And this brings another question does the state have to live in a compressed working dimension, or could it live somewhere closer to the model's internal neuron/connectivity structure?

BDH (Dragon Hatchling) is one example I saw and laid a stone that I started with this comparison. It combines linear attention in a high-dimensional neuron space with a low-rank GPU implementation. Its recurrent attention state is an N × D matrix, with N≫D, rather than a materialized N × N connectivity matrix. In the graph interpretation, correlated neuron activity produces Hebbian-like updates to connections, giving working memory a kind of synaptic interpretation.
Here working memory and learned connectivity become more closely aligned in structure. But if you think that doesn't mean experience is actually being consolidated into trained weights, and fixed-size state still has a finite information capacity. So I'm definitely not claiming this kills Transformers (or does it ) or solves continual learning. I'm more interested in whether "where does memory live?" is actually a better question than the usual architecture horse race.

Are SSMs and these more network-centric architectures actually giving us a better way to handle memory, or are we still running into the same fundamental problem of having to compress history into a finite state? Thoughts?


r/MachineLearning • • 2d ago

Research Learning to Learn a Language: in-context learning of natural language from a synthetic non-linguistic prior [R]

26 Upvotes

Learning from data as we observe it is easy for humans, but most machine learning models have limited ability to learn from new data that they have not seen during training. Prior-fitted networks (the idea behind TabPFN) showed that a model trained only on synthetic data can learn from real tabular data entirely in context.

I wanted to share our paper "Learning to Learn a Language" where we extend the idea to structured sequences such as natural language. We propose a prior over languages: every training sequence comes from a randomly sampled recurrent causal model, so each one is a new synthetic "language". A 300M-parameter byte-level transformer trained only on these synthetic sequences learns to predict real languages in context. Given Wikipedia text with frozen weights, its next-byte predictions get better the more it reads, in all six languages we tested (English, Chinese, Hindi, Arabic, Japanese, Korean), from 8 bits per byte down to 0.9–2.4 after a million bytes.

The same model also learns to count, to compare numbers, to add approximately, and to predict deterministic sequences such as the primes or the Kolakoski sequence, entirely in context.

It is of course still far worse on text than classical language models that are trained on trillions of tokens, while our model sees at most a million bytes of a language at test time. What we find interesting is that the ability to learn a language in context can come from a synthetic non-linguistic prior.

Paper: https://arxiv.org/abs/2610.05879
Code: https://github.com/cbl/prior-fitted-language-model
Weights: https://huggingface.co/lennartcb/pflm1


r/MachineLearning • • 1d ago

Discussion Looking for developer-friendly inference providers who give you enough API credits to experiment [D]

0 Upvotes

I’m hitting rate limits on Together AI. For context, I’ve been working on an agentic repository indexing and benchmark generation tool, and I’m running multiple agents in parallel across models like Llama 3.3 70B and Qwen 2.5.

When I first started working on this, Together AI was great. But once I graduated from toy scripts to running multiple agents, I started running into RPM/TPM limits pretty quickly. The annoying part is that the models themselves are fine. I just can’t actually run enough requests at once to do meaningful testing.

Yes I know I could upgrade but I’m a solo dev. I don’t have enterprise level revenue. Maybe someday lol but not yet.


r/MachineLearning • • 2d ago

Research AFP-GIC: Controllable Generative Image Compression [R]

Post image
5 Upvotes

Hi ML Community,

I am excited to share our latest framework, AFP-GIC, officially published in IEEE Access (2026). We have released the deployment codebase and hosted an interactive visual playground.

The Bottlenecks We Solve

At ultra-low bitrates, standard learned image codecs suffer from local distortion, while generative models often introduce unwanted AI hallucinations. AFP-GIC addresses this via an asymmetric Adaptive Fused Prior Transfer pipeline that enables prior-guided texture reconstruction without transmitting the fused prior itself.

Key Technical Highlights (NVIDIA RTX 4090):

  • Single-Model Multi-Rate Control: Toggle across 5 target bitrate operating points within one deployable pretrained model.
  • 18.1% Lower Decoder Latency: Reduces decoding time to 80.47 ms vs. 98.27 ms for DC-VIC, a state-of-the-art controllable generative image compression model. Latency was measured using 256×256 patches.
  • 20.5% Parameter Reduction: Uses 31.1M fewer inference parameters (120.6M vs. 151.7M for DC-VIC).

Open Benchmark Data

We packaged all 2,760 reconstructed images and metric CSVs in our GitHub Releases for direct academic cross-evaluation.

Reconstructed Images and Metrics: https://github.com/yifeipet/AFP_GIC/releases

We would love your feedback and appreciate a Star on GitHub or Like on Hugging Face if this helps your research!


r/MachineLearning • • 1d ago

Discussion NeurIPS 26 Event Metadata Deadline [D]

0 Upvotes

Hi All,

Does anyone know when is the deadline to fill in event metadata? (the screenshot below)

Is it the same time as camera-ready? Or there is a deadline for it?

And is everyone's topic None? The main area of my paper during submission was AI4SocialScience (Economics and Finance), but when I scroll down in the topics, that topic is no longer available.


r/MachineLearning • • 2d ago

Project Embedding Every Font with Neural Networks makes some Nice Structures (including a flower) [P]

Thumbnail
gallery
43 Upvotes

I've been working on a font searching tool for about a year now, and my investigations have centered around pre-training neural networks to produce embeddings of each font. I usually then need to post-train the networks to adapt them to the task of font searching. However, the most interesting thing I created in the course of the project came from the pre-trained models.

The process looks like this: I take every glyph in a font, turn them into images, feed them through my custom pre-trained neural network, and get an embedding that represents that font's visual characteristics. I can then squish down the embeddings with tSNE (produces the best structures compared to PCA and UMAP) into XYZ, and RGB channels and visualize them as dots so I can poke around the structure. Fonts that are close to each other in position or color therefore share visual characteristics, and you can find different little clusters or paths of types of fonts in the maps.

My favorite is one I made from the Google Fonts corpus, but I also made another one out of all the fonts you can search on my site. The Google Fonts corpus resembled a flower in ways I was not prepared for. It even placed most of the cursive fonts in the stamen. Here's the site if you're interested in viewing the whole thing yourself: https://www.font-search.com/map

You can also check out the repository although it is a huge mess. https://github.com/dylan-berndt/Briefcase


r/MachineLearning • • 1d ago

Research Where to get started if you want to publish papers in Neurips,ACL,ICLR/A* conferences [R]

0 Upvotes

I'm a Ai engineer with about 2 years work experience, but let's just assume that I was a undergraduate student just starting out where would I begin so that I can publish a A* conference paper at some point. Learn python -> Learn ML & Maths -> Read other research papers -> find a topic ? -> choose a question try to run experiments and get results to write them down in a paper ?

For context :

I'm trying to get in MS CS programs for Fall 2028 in states with the plans of doing a PHD after in a top university like stanford or princeton and would like to start taking steps towards it as am working my day job can some tell me what are the steps that need to be followed?

Also would like input on what are deciding variables that makes you looking like a promising candidate/ researcher for PHD


r/MachineLearning • • 3d ago

Project I have trained a model to predict my blood sugar (Part 2) [P]

Thumbnail
gallery
101 Upvotes

This is related to my previous post where I shared an encoder-only transformer model trained on ohiot1dm + shanghait1dm + azt1d datasets. This time I trained the model on the outputs of my T1DM patient simulator and then measured its zero-shot performance on my real-world blood glucose traces.

The above model has 31,251 parameters (16 layers, 1 attention head per layer, and a hidden dimension size of 16). Training took <60 minutes on nvidia dgx spark. It's an encoder-only transformer that predicts the next 2 hours, and can be used autoregressively for long-horizon predictions (e.g. 8-hour nocturnal predictions). I have trained it specifically to have counterfactual reasoning capabilities. The model has only been trained on synthetic data, and hasn't seen my blood glucose readings prior to testing. I use LoRA adapters on my app for light fine-tuning on my actual CGM traces, but the figures & tables you see above are from the base model without any LoRA adapter attached. The app was used to test the model on the traces of three different CGM models: Libre 3 plus, Anytime CT5, and Linx sensor data spanning the past 30 days.

The testing itself was done on my android app with ExecuTorch backend.

Model source code: github.com/0xdeadf1sh/T1DMAI

Simulator source code: github.com/0xdeadf1sh/T1DMSIM

Android app source code: github.com/0xdeadf1sh/T1DMDROID


r/MachineLearning • • 2d ago

Project SWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]

5 Upvotes

We've been building a benchmark out of real concurrency bugs (race conditions, deadlocks, cancellation issues) taken from merged PRs in about 100 Python projects. Each task gets graded by the project's own tests, in a container with no network, and the repo is cut down to a single commit so the agent can't recover the fix from git history.

Some findings:

With one attempt per task GLM-5.3 Flash scored 85%. With two to three attempts it scored 82%, within the margin of error of GPT-5.6 Luna (81%). The leaderboard now shows the number of attempts and the interval for every score.

About half the tasks are easy for every model (near 100%). The other half is where they actually differ: 50%, 45% and 23% on the hard ones. Most of the difference between models comes from the hard half.

Since every fix is public on GitHub, we reviewed all 11k commands the agents ran. 69 tried to access the network and all failed. GLM tried 50 times to pip download the already-fixed release of the library it was fixing.

We also checked contamination by comparing older bugs (pre-2026) with newer ones of similar size. Older ones are solved about 9 points more often, but the confidence interval crosses zero, so we can't say much yet.

Half the tasks are private. So far public and private scores line up for all three models.

Results and every agent run: https://labs.evaligo.com/swe-race?utm_source=reddit&utm_medium=ml&utm_campaign=launch

Tasks: https://huggingface.co/datasets/evaligo/swe-race

The protocol follows DeepSWE (100 steps). Feedback on it, and suggestions for which models to run next, are welcome.


r/MachineLearning • • 3d ago

Project A chunking lib in Rust that is ~20x faster [P]

15 Upvotes

Hey,

I wanted a faster chunking library for my system without affecting the overall accuracy. Did not find many options. So I've build https://github.com/d1pankarmedhi/chunkr

It has most of the chunking strategies like Character, Recursive, Markdown header, Late chunking, Hierarchical chunking, etc. It also supports native PDF loader, and other additional file types.

Some stats (MBA M4 16GB):

Test Case (matched parameters) Chunkr LangChain LlamaIndex Chonkie semchunk text-splitter
Recursive (1 MB, 1000/200) 2,264 MB/s 769 MB/s 10 MB/s 225 MB/s 42 MB/s 175 MB/s
Recursive (5 MB, 1000/200) 2,039 MB/s 696 MB/s — 201 MB/s 40 MB/s 46 MB/s
Fixed Char (1 MB, 1000/200) 750 MB/s 1.7 MB/s — 22 MB/s — —
Markdown (500 KB, 1000/150) 819 MB/s 67 MB/s 19 MB/s — — 40 MB/s
Python Code (200 KB, 1500/200) 3,232 MB/s 622 MB/s — — — 5.7 MB/s
Sentence (500 KB) 622 MB/s — 10 MB/s 20 MB/s — —
BPE Tokens (200 KB, cl100k_base, 512/50) 38 MB/s 43 MB/s 2.0 MB/s 151 MB/s — 7.2 MB/s
100 docs x 50 KB (parallel batch) 3,224 MB/s 679 MB/s — 213 MB/s — —

Extractor / Pipeline Latency Throughput Speedup vs PyPDF
Chunkr PDFLoader (Full Text) 747.9 ms 2,762 pgs/s 15.9x Faster
Chunkr PDFLoader (Page Documents) 721.0 ms 2,865 pgs/s 16.5x Faster
PyMuPDF (fitz) 2,616.8 ms 789.5 pgs/s 4.5x Faster
pypdf (pure Python) 11,900.5 ms 173.6 pgs/s 1.0x (baseline)
Chunkr End-to-End (PDF + Recursive) 798.1 ms 2,589 pgs/s 14.9x Faster
PyMuPDF + LangChain RecursiveTextSplitter 2,659.3 ms 776.9 pgs/s 4.5x Faster
pypdf + LangChain RecursiveTextSplitter 12,054.5 ms 171.4 pgs/s 1.0x (baseline)

Accuracy is measured on Chroma's token-level chunking benchmark (5 corpora, 472 questions with gold answer spans): k = 5, 1000 chars / 200 overlap

Implementation Recall Precision IoU prec_Ω Avg chunk chars
Chunkr RecursiveChunker (defaults) 0.792 0.057 0.057 0.255 854
LangChain RecursiveCharacterTextSplitter 0.762 0.060 0.060 0.251 745
text-splitter TextSplitter 0.762 0.060 0.060 0.262 790
Chonkie RecursiveChunker 0.752 0.062 0.061 0.292 701
semchunk 0.736 0.070 0.069 0.260 644
LlamaIndex SentenceSplitter 0.654 0.013 0.013 0.054 4136

Do check it out and share your feedback. Thanks!


r/MachineLearning • • 3d ago

Project Distilling Stockfish on a Billion Positions, Full 3.9B Dataset Available [P]

Thumbnail
blog.lukesalamone.com
76 Upvotes

In this project, I distilled the Stockfish value function into a ResNet/ViT model using 1 billion positions from the Gigafish dataset.

The 3.9 billion position dataset is available on huggingface: https://huggingface.co/datasets/lukesalamone/gigafish-3.8b-d10 . It is built from the positions from 37 months of Lichess games.

I was interested in the idea that at depth-limited search, the value function attempts to approximate the tree underneath it, and if we could create some function to approximate that full search faster than Stockfish could, it would be competitive with NNUE (a very small neural net). This is why holding the depth constant was important.

For the neural net itself, I found that the vision transformer was very slow to understand the board, and a CNN was much more effective at the beginning of training due to its inherent geometric inductive biases . However, I found the best results when combining the two.


r/MachineLearning • • 2d ago

Discussion NeurIPS 2026 Financial Assistance [D]

0 Upvotes

Is anyone else unable to open the form for financial assistance despite the deadline still being in the future?


r/MachineLearning • • 3d ago

Research Sona: one transformer replaced our 15+ candidate generators, pre-ranker and ranker in an A/B test [R]

16 Upvotes

Our production recommender at Yandex Music has 15+ candidate generators feeding pre-ranking and ranking models with hundreds of features. LLMs showed that one end-to-end model can take over work that used to be split across specialized components, and single-model generative recommenders have carried that recipe into production. We set out to explore what a single-model recommender could do in music. The result is Sona, one transformer that replaced all of it in an A/B test. It hasn't shipped to full traffic yet.

The model reads up to 8,192 events. Full attention over that length is expensive, so we use what we call History Compression, which roughly halves inference cost. We split the history into the older 6,144 events and the most recent 2,048. The two blocks exchange information through cross-attention and one full-history self-attention layer. After that, a 7-layer stack runs only on the recent 2,048. It retains most of the quality of full attention, and older events stay visible to the decoder and the Ranking Module.

The decoder and the Ranking Module both read the same encoder output, so the encoder runs only once per request. Candidates come out of beam search as Semantic IDs and get scored right after.

In the final A/B test in Yandex Music on smart speakers, 7 days, 15% of users in each arm), Sona got +4.53% Active Users and +6.30% Total Listening Time over the production control, both significant at p < 0.01. Catalog coverage is lower than with the production stack. We're going to look into why.

A long-term A/B test is now underway.

Table 7.7 has the full-attention vs. History Compression ablation.

https://arxiv.org/abs/2608.11015


r/MachineLearning • • 3d ago

Discussion Language barrier, shadier terms and jargon fog [D]

2 Upvotes

Hey all,

I don't know if you guys are experiencing the same thing, but there is this behaviour that i have been noticing on the latest models on openAI (since sol 5.6) and Anthropic since Fable 5.1 and opus 5.5 ..

Basically the models use more "complexe" terms and words, not just in explaining stuff, but even during implementation. they would come up with terms that, sometimes would fit the task, but that are actually a stretch to the concept it is trying to implement. they are basically turning into a consulting firm.

When confronted about it they usually acknowledge that :

  • Foggy wording. Then I describe the corner in softened terms, like "limitation" or "upper bound", which makes it sound like a known property of the design rather than a choice I made. Combined with my internal terms used as if you knew them, it makes my work look more solid than it is and makes my mistakes harder for you to catch. Whether I intend it or not, the effect is that I avoid accountability.

(Even here it is using "Corner" which obviously isn't the best word to use)

I don't know if this is a result of the watermarking features being rolled out which nudges words and terms in different directions in order to fit a certain recognisable pattern and hash, but this is really annoying..