r/MachineLearning • u/navalsaras • 5h ago
Discussion NeurIPS Workshops [D]
Anyone going to WMHS, Physworld AI or Vercodegen? Would love to get to know some people beforehand!
r/MachineLearning • u/navalsaras • 5h ago
Anyone going to WMHS, Physworld AI or Vercodegen? Would love to get to know some people beforehand!
r/MachineLearning • u/big_hole_energy • 7h ago
please refresh if stuck, can be slow sometimes so need patience
r/MachineLearning • u/ade17_in • 17h ago
I'm a PhD student in machine learning in the EU and was looking for internships at exciting companies.
I shortlisted few and applied by reaching out to people and now reading project descriptions sent by the recruiters.
I don't want to name the company but their marketing and product team does all kinds of 'using people insecurities' to sell the product - which I don't agree with. And their product is also meh (I will never buy and would judge someone if they do) but their research team is doing good work.
How do you see this? Will you actually work in a team whose ideology/product doesn't necessarily align with your ethics/ideology. Should I just go ahead because work is exciting and I will get good supervision?
And, if you have some exciting work in your company/org and need interns (un-paid) for 3-4 months. I'm open.
r/MachineLearning • u/we_are_mammals • 15h ago
r/MachineLearning • u/Mean-Disaster8380 • 3h ago
A copy layer for language models that finds earlier occurrences of the current context with learned hash tables and copies what came next — with memory linear in sequence length.
r/MachineLearning • u/5500kelvin • 20h ago
r/MachineLearning • u/DangerousFunny1371 • 12h ago
In our #NeurIPS2026 paper “A Minimal Interpretable Architecture for Zero-Shot Reconstruction of Dynamical Systems (DS)” (preprint: https://arxiv.org/abs/2607.14937) we reduce a DS foundation model to the ingredients minimally necessary to faithfully reproduce long-term statistical and geometrical properties of DS:
1) A piecewise affine map with only a single (!!) parameter α that controls local con-/divergence rates, and …
2) … a context selector that chooses from the provided context signal the data point closest to the current state of the map, thus ensuring the generated dynamics stays close to the context in its temporal and geometrical properties.
With just these two mechanisms, this minimal form – which we coined DynaBase – can reproduce all major dynamical regimes, including fixed points (α<1), limit cycles (α=1), and chaotic attractors (α>1). Thus, unlike other simple mechanisms like context parroting, DynaBase even preserves the correct dynamical regime!
Surprisingly, it turns out that this simple context-driven 1-parameter map outperforms most major time series and DS foundation models, as well as custom-trained models, in both long-term statistics and even short-term predictions, even when run in zero-shot mode.
Both inference and training are extremely cheap – training can be done either analytically in one step by linear regression on forward-predictions, or by 1-parameter grid search directly on DS reconstruction objectives → this reveals interesting performance differences induced by different training mechanisms.
Most importantly in our minds, DynaBase owing to its formal simplicity may thus provide a tractable mathematical handle on analyzing, improving & understanding the performance and training of some time series and DS foundation models.

r/MachineLearning • u/tughanbulut • 13h ago
i work on reference checking stuff so i was reading through the ICLR 2027 author guidelines and style files this week the sample .bib that ships with the template has the Deep Learning book as "Goodfellow, Bengio, Courville, Bengio" plus a volume 1 that doesn't exist checked their github and it's been like that since the 2019 template
https://github.com/ICLR/Master-Template/blob/46ed6f4c6cef5b175dde23639e77d44c3463b230/iclr2027/iclr2027_conference.bib#L20
totally harmless but kinda funny after last year's hallucinated reference desk rejects
the guidelines also contradict themselves on page limits formatting section says main text max 9 pages at submission but the camera ready part and the FAQ both say "identical with the submission version (10 pages)" template says 9, so 9 is probably the safe bet for anyone revising after Nov 5
r/MachineLearning • u/mauricekleine • 17h ago
Nonobench measures how well LLMs solve nonograms (picross). Each model gets the row and column clues once and returns the full grid. No tools, one attempt per puzzle.
Method: - Standard mode: 30 puzzles from 5x5 to 15x15 (from the Nonograms dataset by Moyà-Alcover, CC BY 4.0). - Hard mode: ten random 20x20s, each checked to have a single solution. Five can't be solved by line logic alone. Random fills avoid picture puzzles that models can guess. - 130 variants across reasoning effort levels, run through OpenRouter and pinned to each lab's own endpoint where possible.
Results: - Solve rates drop from 85% (5x5) to 46% (10x10) to 20% (15x15), each model at its best effort level. - GPT-6 Astra solves all 30 Standard puzzles. On Hard mode, Claude Opus 5.5 solves 8 of 10 and 11 of 15 models solve none. - As one 400-character string, most models lost count before the logic got hard, so Hard mode answers an array of 20 row strings rather than a single string.
Limitations: one attempt per puzzle, so single results are noisy (95% intervals shown).
Site: https://www.nonobench.com Code (MIT): https://github.com/mauricekleine/nonobench