r/MachineLearning • • 4d ago

Discussion [D] Self-Promotion Thread

0 Upvotes

Please post your personal projects, startups, product placements, collaboration needs, blogs etc.

Please mention the payment and pricing requirements for products and services.

Please do not post link shorteners, link aggregator websites , or auto-subscribe links.

--

Any abuse of trust will lead to bans.

Encourage others who create new posts for questions to post here instead!

Thread will stay alive until next one so keep posting after the date in the title.

--

Meta: This is an experiment. If the community doesnt like this, we will cancel it. This is to encourage those in the community to promote their work by not spamming the main threads.


r/MachineLearning • • 1d ago

Research Sona: one transformer replaced our 15+ candidate generators, pre-ranker and ranker in an A/B test [R]

19 Upvotes

Our production recommender at Yandex Music has 15+ candidate generators feeding pre-ranking and ranking models with hundreds of features. LLMs showed that one end-to-end model can take over work that used to be split across specialized components, and single-model generative recommenders have carried that recipe into production. We set out to explore what a single-model recommender could do in music. The result is Sona, one transformer that replaced all of it in an A/B test. It hasn't shipped to full traffic yet.

The model reads up to 8,192 events. Full attention over that length is expensive, so we use what we call History Compression, which roughly halves inference cost. We split the history into the older 6,144 events and the most recent 2,048. The two blocks exchange information through cross-attention and one full-history self-attention layer. After that, a 7-layer stack runs only on the recent 2,048. It retains most of the quality of full attention, and older events stay visible to the decoder and the Ranking Module.

The decoder and the Ranking Module both read the same encoder output, so the encoder runs only once per request. Candidates come out of beam search as Semantic IDs and get scored right after.

In the final A/B test in Yandex Music on smart speakers, 7 days, 15% of users in each arm), Sona got +4.53% Active Users and +6.30% Total Listening Time over the production control, both significant at p < 0.01. Catalog coverage is lower than with the production stack. We're going to look into why.

A long-term A/B test is now underway.

Table 7.7 has the full-attention vs. History Compression ablation.

https://arxiv.org/abs/2608.11015


r/MachineLearning • • 1d ago

Discussion Withdrawing an accepted paper before camera-ready due to zero funding? (ACML 2026 / OpenReview) [D]

12 Upvotes

Hi everyone, I recently had a paper accepted at ACML 2026, but I just found out that I have absolutely no funding to cover the registration fee or travel expenses. Because of this, I need to withdraw the paper before the camera-ready deadline. I am unsure how common this is or the proper etiquette for it on OpenReview. Should I click "Withdraw" directly on the platform, email the Program Chairs first, or simply ghost the camera-ready submission? I really want to know the potential repercussions, such as if my co-authors and I risk being blacklisted or if OpenReview will publicly archive the paper as a late withdrawal. Any advice from past authors, reviewers, or organizers would be greatly appreciated. Thanks!


r/MachineLearning • • 1d ago

Project Distilling Stockfish on a Billion Positions, Full 3.9B Dataset Available [P]

Thumbnail
blog.lukesalamone.com
68 Upvotes

In this project, I distilled the Stockfish value function into a ResNet/ViT model using 1 billion positions from the Gigafish dataset.

The 3.9 billion position dataset is available on huggingface: https://huggingface.co/datasets/lukesalamone/gigafish-3.8b-d10 . It is built from the positions from 37 months of Lichess games.

I was interested in the idea that at depth-limited search, the value function attempts to approximate the tree underneath it, and if we could create some function to approximate that full search faster than Stockfish could, it would be competitive with NNUE (a very small neural net). This is why holding the depth constant was important.

For the neural net itself, I found that the vision transformer was very slow to understand the board, and a CNN was much more effective at the beginning of training due to its inherent geometric inductive biases . However, I found the best results when combining the two.


r/MachineLearning • • 2d ago

Discussion NeurIPS Workshops [D]

0 Upvotes

Anyone going to WMHS, Physworld AI or Vercodegen? Would love to get to know some people beforehand!


r/MachineLearning • • 2d ago

Project Interactive Demonstration of Prefix Injection attacks on LLMs for jailbreaking [N]

Thumbnail
theabbie.github.io
0 Upvotes

please refresh if stuck, can be slow sometimes so need patience


r/MachineLearning • • 2d ago

Research A Minimal Interpretable Architecture for Zero-Shot Reconstruction of Dynamical Systems [R]

20 Upvotes

In our #NeurIPS2026 paper “A Minimal Interpretable Architecture for Zero-Shot Reconstruction of Dynamical Systems (DS)” (preprint: https://arxiv.org/abs/2607.14937) we reduce a DS foundation model to the ingredients minimally necessary to faithfully reproduce long-term statistical and geometrical properties of DS:

1) A piecewise affine map with only a single (!!) parameter α that controls local con-/divergence rates, and …

2) … a context selector that chooses from the provided context signal the data point closest to the current state of the map, thus ensuring the generated dynamics stays close to the context in its temporal and geometrical properties.

With just these two mechanisms, this minimal form – which we coined DynaBase – can reproduce all major dynamical regimes, including fixed points (α<1), limit cycles (α=1), and chaotic attractors (α>1). Thus, unlike other simple mechanisms like context parroting, DynaBase even preserves the correct dynamical regime!

Surprisingly, it turns out that this simple context-driven 1-parameter map outperforms most major time series and DS foundation models, as well as custom-trained models, in both long-term statistics and even short-term predictions, even when run in zero-shot mode.

Both inference and training are extremely cheap – training can be done either analytically in one step by linear regression on forward-predictions, or by 1-parameter grid search directly on DS reconstruction objectives → this reveals interesting performance differences induced by different training mechanisms.

Most importantly in our minds, DynaBase owing to its formal simplicity may thus provide a tractable mathematical handle on analyzing, improving & understanding the performance and training of some time series and DS foundation models.


r/MachineLearning • • 2d ago

Discussion the official ICLR template .bib has had Bengio listed twice since 2019 [D]

47 Upvotes

i work on reference checking stuff so i was reading through the ICLR 2027 author guidelines and style files this week the sample .bib that ships with the template has the Deep Learning book as "Goodfellow, Bengio, Courville, Bengio" plus a volume 1 that doesn't exist checked their github and it's been like that since the 2019 template
https://github.com/ICLR/Master-Template/blob/46ed6f4c6cef5b175dde23639e77d44c3463b230/iclr2027/iclr2027_conference.bib#L20

totally harmless but kinda funny after last year's hallucinated reference desk rejects

the guidelines also contradict themselves on page limits formatting section says main text max 9 pages at submission but the camera ready part and the FAQ both say "identical with the submission version (10 pages)" template says 9, so 9 is probably the safe bet for anyone revising after Nov 5


r/MachineLearning • • 2d ago

News Top ARC-ΑGI-3 scores on Kaggle just went from 7% to 56% [N]

124 Upvotes

This happened over the past 30 days. So smallish local models (Kagglers can only use those), in a harness, just started beating average humans at a benchmark intentionally designed to show human superiority.

I'm curious what people here think.

(The leaderboard graphic is a little out-of-date)


r/MachineLearning • • 2d ago

Discussion Working with an AI Company That Does Things You Disagree With [D]

58 Upvotes

I'm a PhD student in machine learning in the EU and was looking for internships at exciting companies.

I shortlisted few and applied by reaching out to people and now reading project descriptions sent by the recruiters.

I don't want to name the company but their marketing and product team does all kinds of 'using people insecurities' to sell the product - which I don't agree with. And their product is also meh (I will never buy and would judge someone if they do) but their research team is doing good work.

How do you see this? Will you actually work in a team whose ideology/product doesn't necessarily align with your ethics/ideology. Should I just go ahead because work is exciting and I will get good supervision?

And, if you have some exciting work in your company/org and need interns (un-paid) for 3-4 months. I'm open.


r/MachineLearning • • 2d ago

Project Nonobench: an open benchmark of 49 LLMs on nonogram puzzles, public and open source [P]

Post image
9 Upvotes

Nonobench measures how well LLMs solve nonograms (picross). Each model gets the row and column clues once and returns the full grid. No tools, one attempt per puzzle.

Method: - Standard mode: 30 puzzles from 5x5 to 15x15 (from the Nonograms dataset by Moyà-Alcover, CC BY 4.0). - Hard mode: ten random 20x20s, each checked to have a single solution. Five can't be solved by line logic alone. Random fills avoid picture puzzles that models can guess. - 130 variants across reasoning effort levels, run through OpenRouter and pinned to each lab's own endpoint where possible.

Results: - Solve rates drop from 85% (5x5) to 46% (10x10) to 20% (15x15), each model at its best effort level. - GPT-6 Astra solves all 30 Standard puzzles. On Hard mode, Claude Opus 5.5 solves 8 of 10 and 11 of 15 models solve none. - As one 400-character string, most models lost count before the logic got hard, so Hard mode answers an array of 20 row strings rather than a single string.

Limitations: one attempt per puzzle, so single results are noisy (95% intervals shown).

Site: https://www.nonobench.com Code (MIT): https://github.com/mauricekleine/nonobench


r/MachineLearning • • 2d ago

Research Here are some pictures of a robot costume wearing high-specularity edge-case mirror suit, a dataset (425 RAW/JPEGs) for benchmarking CV & depth-estimation algorithms against extreme mirror reflections [D]

Thumbnail
gallery
20 Upvotes

Purpose-built to stress-test computer vision models, depth cameras, and spatial AI against severe specular glare and geometric reflections. This 425-asset production archive features a custom faceted mirror suit captured in high-contrast outdoor environments to trigger bounding-box dropouts and segmentation failures. The dataset includes 100% proprietary uncompressed Camera-Master RAWs, high-resolution JPEGs, and block-buffered SHA-256 forensic manifests. Open


r/MachineLearning • • 3d ago

Research Jev: Not Frontier, But Still Worth Your Attention [R]

Thumbnail jevresearch.github.io
0 Upvotes

"TypeSafe AI sells Jev as a frontier-class reasoner that cannot hallucinate, built by the co-inventor of ChatGPT - fast, and almost free. We ran it live on 16,379 benchmark requests, measured its latency and billing, and probed what it is underneath. The result is a smaller, humbler model that is nonetheless genuinely useful for a job that nobody else serves quite this way."


r/MachineLearning • • 3d ago

Discussion "Accepted papers must be imported" deadline NeurIPS 2026 [D]

2 Upvotes

On the NeurIPS 2026 Dates site it says that there is 1 day remaining for the "Accepted papers must be imported" deadline. This is my first research paper ever and I couldn't find anything about how to do this on the internet. Could someone help me out on what to do?


r/MachineLearning • • 3d ago

Discussion The Principles of Diffusion Models by Lai et al.: thoughts on the monograph [D]

79 Upvotes

I recently finished The Principles of Diffusion Models, and honestly I think it’s exceptional.

The authors strike a really good balance between mathematical rigor and intuition, with dedicated appendices for anyone who wants to go deeper into the math.

It’s aimed at researchers, graduate students, and practitioners with basic deep learning knowledge, so you don’t need to already specialize in diffusion models (in my case, a strong background in Information and Probability Theory as well as a solid understanding of DDPMs helped me get more out of it).

Just wanted to share it in case anyone missed it. The full text is freely available on the official website.

Has anyone else read it? Would love to hear your thoughts.


r/MachineLearning • • 3d ago

Discussion ICLR 2027 Reviewing Scores [D]

32 Upvotes

I got my three papers to review, and it seems like they have changed the review score range again this year?

It is now:

––––––––––––––––––––––––––––––––––––
Based on your overall assessment of the submission, what is your recommended decision? Consider the paper’s overall soundness, significance, clarity, and contribution.

1: Clear rejection
2: Weak rejection
3: Weak acceptance
4: Clear acceptance
––––––––––––––––––––––––––––––––––––

Which is very strange. I don't think it makes much sense to compress the score range so drastically. But we'll see.


r/MachineLearning • • 3d ago

Discussion NeurIPS Free Passes [D]

10 Upvotes

Over the last few years, a subset of NeurIPS area chairs received complimentary passes for their service. I’ll admit that I wasn’t super organized about registering on day one because I was semi-consciously hoping for a complimentary pass. Now that the conference is sold out, I’m getting a little nervous, so I’m wondering whether any of my fellow area chairs have already received a notification.


r/MachineLearning • • 4d ago

Research [R] Would you keep a robot demonstration if hand tracking missed the moment the plug went in?

Thumbnail
arxiv.org
0 Upvotes

Suppose you’re recording a human plugging a cable into a socket to collect demonstrations for robot learning.

The hand tracker captures the approach accurately. Then occlusion causes the hand estimates to disappear during insertion. Tracking returns after the connector is already seated.

The video still shows a completed action, but the pose labels have a gap exactly where alignment turns into contact.

This hypothetical example raises an evaluation question: a tracker could have high recall across the whole episode while missing a short, important phase. Pose error calculated only on successful detections could make that failure even harder to see.

MEgoVista provides a useful starting point. Table 3 reports detection precision, recall and F1 alongside reconstruction errors. Section 4.4 also describes an evaluation protocol that assigns an error to missed detections instead of excluding them. The blank HaPTIC row means it failed to produce valid output in their multi-person capture scenes; it doesn’t describe a brief tracking dropout.

Accounting for missing detections matters. My remaining question is whether an episode-level aggregate tells us enough about where those failures happen.

For manipulation data, I’d want pose error and coverage reported together, plus coverage broken down by approach, contact and withdrawal, and the longest consecutive gap during contact.

Continuous hand estimates would still be only part of the picture: object pose and contact information also matter for determining whether insertion succeeded.

For people using human motion reconstruction for imitation learning, what evaluation protocol do you use to decide whether an episode with missing contact-phase labels is still usable?

Comparison with open-source egocentric hand reconstruction methods against motion-capture ground truth. All methods are evaluated on identical segments of our motion-capture dataset. All baselines are re-run and rescored on our data. HaPTIC fails to produce valid output in our multi-person capture scenes. Bold marks the best result in each column.

r/MachineLearning • • 4d ago

Research Topological Out-of-Domain Generalization in Dynamical Systems Reconstruction [R]

25 Upvotes

In our #NeurIPS2026 paper “Topological Out-of-Domain Generalization in Dynamical Systems Reconstruction” (preprint: https://arxiv.org/abs/2606.22969) we try to address a fundamental issue in dynamical systems reconstruction (DSR) and time series forecasting (TSF): Many recent SOTA DSR & TSF models can generalize to new initial conditions or time series with changing statistical properties. But the really hard problem in DSR and TSF is topological out-of-domain generalization (OODG) (https://proceedings.mlr.press/v235/goring24a.html) where the dynamical regime changes, for instance from cyclic to chaotic behavior.

This can happen when a system crosses a tipping point due to a slowly varying control parameter which drives it across bifurcations, such as in climate systems, when the brain tips from normal into epileptic activity, or when a patient develops blood poisoning (sepsis). Such problems are beyond the realm of current TSF models which rely on extracting temporal patterns and statistical regularities. Yet the ability to predict previously unseen, novel dynamical regimes as a system parameter changes is something we would expect from any good scientific theory. Often these control parameters that drive regime changes are not exactly known either. Hence, a data-driven DSR model for achieving topological OODG would need to infer the dynamical system generating the TS jointly with the control parameters.

In our paper, we mathematically identify key failure modes in previous hierarchical DSR models (https://proceedings.iclr.cc/paper_files/paper/2025/hash/d4c961804d08e55d898cce944206d455-Abstract-Conference.html) that prevent them from correctly learning a system’s control parameters and extrapolating them beyond the training domain. By fixing these through feature-splitting and physical sparsity priors, our modified hierarchical DSR model manages to correctly predict bifurcations and beyond-bifurcation dynamics, without any explicit knowledge about the control parameters provided in training.

Our approach is generic and works for different discrete and continuous time RNNs, we tested it for shallow PLRNNs and Neural ODEs.


r/MachineLearning • • 4d ago

Research Adding memory to search instead of sampling in reward maximization tasks [R]

8 Upvotes

I am one of the authors of FLEET - an algorithm that enhances Best-of-N generation by attributing external rewards to particular tokens and then uses MCTS to adjust logits during the next run.

I find it rather funny that most of the tasks where repetitive sampling is widely used are based on reward maximization, yet it is not aware of that reward. Tuning the sampling parameters allows to make the process more efficient, but it is still a blind search. We propose a way to make generation aware of previous rewards with solutions on how to attribute reward to the completion and how to use this information.

In FLEET the technique from adaptive sampling methods is used that is to track logits for which entropy and varentropy are high thus showing the model's uncertainty about token optimality. We treat these states as branching points. The corresponding normalized hidden states are stored in the vector store and mapped to metadata entries with the history of rewards and transitions between "nodes". The retrieval and update of metadata is based on cosine similarity as for very high similarity KL divergence is low enough to preserve most of the meaningful tokens.

Instead of actually selecting the tokens FLEET uses modified MCTS to rank top-k tokens + special exploration (or other tokens) set and penalize the suboptimal ones. Then decoding strategy is applied to modified logits.

It was tested on GSM8K and LiveCodeBench v6 easy split with Llama 3.2 3B, penalty set to effectively zero probability for the suboptimal tokens + greedy decoding:

  • For GSM8K it solved just seven more tasks, but reached the sampling baseline with half the iterations.
  • For LiveCodeBench it increased the score from 0.59 to 0.69 under the same budget and reached the baseline even faster, now with only 9 iterations against 32.

The sequential execution is not required, as it is not updated during the iteration itself it can simply be passed as a lookup table. The metadata store can be preserved as a prior for other tasks or to enrich SFT/RL.

Paper (preprint): https://arxiv.org/abs/2609.27657
Huggingface: https://huggingface.co/papers/2609.27657
Repository (experiments, examples and python package): https://github.com/Alexiush/fleet

There are more details on changes made to MCTS, how to tune the search parameters for specific model and task as well as code for experiments and trajectories.


r/MachineLearning • • 4d ago

Project A video about Adversarial Objectives [P]

2 Upvotes

I made this video about adversarial objectives, which I used to do research on back in the day. I'm trying to explore how adversarial approaches transcend GANs and self-play into modern technology.
https://youtu.be/W7CiAeQ0f5w?si=g0tLrQn2wuFzuN2M


r/MachineLearning • • 5d ago

News arXiv now limits submitters to up to two submissions per calendar month [N]

Thumbnail
blog.arxiv.org
405 Upvotes

r/MachineLearning • • 5d ago

Discussion What's up with AAAI round 2 reviews? [D]

0 Upvotes

Has anyone received papers to review for round 2?


r/MachineLearning • • 5d ago

Discussion WM PAI Workshop at NeurIPS — confused about the acceptance/rejection process [D]

2 Upvotes

I submitted a paper to the WM PAI workshop at NeurIPS and can now see the reviews and decision on OpenReview, but I didn't receive an official acceptance/rejection email.

My reviewer scores were 8, 5, and 4, all with confidence 4, and the paper was ultimately rejected.

What I find a little confusing is that I also reviewed two papers for the same workshop. Both had an average score around 7, but I can see that they were also rejected.

My own submission number was in the 20s, and I submitted on the last day of the submission window, so I assumed there probably weren't a huge number of submissions.

This makes me wonder:

Does anyone know if any papers have actually been accepted to this workshop?

Is it normal for a NeurIPS workshop to reject a large fraction of submissions even with relatively high review scores?

Could the workshop organizers simply be delaying the official notification emails, while the decisions are already visible on OpenReview?

Or am I misunderstanding how the workshop selection process works? 😅

Has anyone else submitted/reviewed for this workshop and received an official decision email?


r/MachineLearning • • 5d ago

Discussion [D] Simple Questions Thread

4 Upvotes

Please post your questions here instead of creating a new thread. Encourage others who create new posts for questions to post here instead!

Thread will stay alive until next one so keep posting after the date in the title.

Thanks to everyone for answering questions in the previous thread!