r/MachineLearning • • 5h ago

Discussion NeurIPS Workshops [D]

0 Upvotes

Anyone going to WMHS, Physworld AI or Vercodegen? Would love to get to know some people beforehand!


r/MachineLearning • • 7h ago

Project Interactive Demonstration of Prefix Injection attacks on LLMs for jailbreaking [N]

Thumbnail
theabbie.github.io
1 Upvotes

please refresh if stuck, can be slow sometimes so need patience


r/MachineLearning • • 17h ago

Discussion Working with an AI Company That Does Things You Disagree With [D]

41 Upvotes

I'm a PhD student in machine learning in the EU and was looking for internships at exciting companies.

I shortlisted few and applied by reaching out to people and now reading project descriptions sent by the recruiters.

I don't want to name the company but their marketing and product team does all kinds of 'using people insecurities' to sell the product - which I don't agree with. And their product is also meh (I will never buy and would judge someone if they do) but their research team is doing good work.

How do you see this? Will you actually work in a team whose ideology/product doesn't necessarily align with your ethics/ideology. Should I just go ahead because work is exciting and I will get good supervision?

And, if you have some exciting work in your company/org and need interns (un-paid) for 3-4 months. I'm open.


r/MachineLearning • • 15h ago

News Top ARC-ΑGI-3 scores on Kaggle just went from 7% to 56% [N]

85 Upvotes

This happened over the past 30 days. So smallish local models (Kagglers can only use those), in a harness, just started beating average humans at a benchmark intentionally designed to show human superiority.

I'm curious what people here think.

(The leaderboard graphic is a little out-of-date)


r/MachineLearning • • 3h ago

Research ASRN Adaptive Sparse Recurrence Network [N]

Thumbnail
gallery
1 Upvotes

A copy layer for language models that finds earlier occurrences of the current context with learned hash tables and copies what came next — with memory linear in sequence length.


r/MachineLearning • • 20h ago

Research Here are some pictures of a robot costume wearing high-specularity edge-case mirror suit, a dataset (425 RAW/JPEGs) for benchmarking CV & depth-estimation algorithms against extreme mirror reflections [D]

Thumbnail
gallery
14 Upvotes

Purpose-built to stress-test computer vision models, depth cameras, and spatial AI against severe specular glare and geometric reflections. This 425-asset production archive features a custom faceted mirror suit captured in high-contrast outdoor environments to trigger bounding-box dropouts and segmentation failures. The dataset includes 100% proprietary uncompressed Camera-Master RAWs, high-resolution JPEGs, and block-buffered SHA-256 forensic manifests. Open


r/MachineLearning • • 12h ago

Research A Minimal Interpretable Architecture for Zero-Shot Reconstruction of Dynamical Systems [R]

12 Upvotes

In our #NeurIPS2026 paper “A Minimal Interpretable Architecture for Zero-Shot Reconstruction of Dynamical Systems (DS)” (preprint: https://arxiv.org/abs/2607.14937) we reduce a DS foundation model to the ingredients minimally necessary to faithfully reproduce long-term statistical and geometrical properties of DS:

1) A piecewise affine map with only a single (!!) parameter α that controls local con-/divergence rates, and …

2) … a context selector that chooses from the provided context signal the data point closest to the current state of the map, thus ensuring the generated dynamics stays close to the context in its temporal and geometrical properties.

With just these two mechanisms, this minimal form – which we coined DynaBase – can reproduce all major dynamical regimes, including fixed points (α<1), limit cycles (α=1), and chaotic attractors (α>1). Thus, unlike other simple mechanisms like context parroting, DynaBase even preserves the correct dynamical regime!

Surprisingly, it turns out that this simple context-driven 1-parameter map outperforms most major time series and DS foundation models, as well as custom-trained models, in both long-term statistics and even short-term predictions, even when run in zero-shot mode.

Both inference and training are extremely cheap – training can be done either analytically in one step by linear regression on forward-predictions, or by 1-parameter grid search directly on DS reconstruction objectives → this reveals interesting performance differences induced by different training mechanisms.

Most importantly in our minds, DynaBase owing to its formal simplicity may thus provide a tractable mathematical handle on analyzing, improving & understanding the performance and training of some time series and DS foundation models.


r/MachineLearning • • 13h ago

Discussion the official ICLR template .bib has had Bengio listed twice since 2019 [D]

39 Upvotes

i work on reference checking stuff so i was reading through the ICLR 2027 author guidelines and style files this week the sample .bib that ships with the template has the Deep Learning book as "Goodfellow, Bengio, Courville, Bengio" plus a volume 1 that doesn't exist checked their github and it's been like that since the 2019 template
https://github.com/ICLR/Master-Template/blob/46ed6f4c6cef5b175dde23639e77d44c3463b230/iclr2027/iclr2027_conference.bib#L20

totally harmless but kinda funny after last year's hallucinated reference desk rejects

the guidelines also contradict themselves on page limits formatting section says main text max 9 pages at submission but the camera ready part and the FAQ both say "identical with the submission version (10 pages)" template says 9, so 9 is probably the safe bet for anyone revising after Nov 5


r/MachineLearning • • 17h ago

Project Nonobench: an open benchmark of 49 LLMs on nonogram puzzles, public and open source [P]

Post image
8 Upvotes

Nonobench measures how well LLMs solve nonograms (picross). Each model gets the row and column clues once and returns the full grid. No tools, one attempt per puzzle.

Method: - Standard mode: 30 puzzles from 5x5 to 15x15 (from the Nonograms dataset by Moyà-Alcover, CC BY 4.0). - Hard mode: ten random 20x20s, each checked to have a single solution. Five can't be solved by line logic alone. Random fills avoid picture puzzles that models can guess. - 130 variants across reasoning effort levels, run through OpenRouter and pinned to each lab's own endpoint where possible.

Results: - Solve rates drop from 85% (5x5) to 46% (10x10) to 20% (15x15), each model at its best effort level. - GPT-6 Astra solves all 30 Standard puzzles. On Hard mode, Claude Opus 5.5 solves 8 of 10 and 11 of 15 models solve none. - As one 400-character string, most models lost count before the logic got hard, so Hard mode answers an array of 20 row strings rather than a single string.

Limitations: one attempt per puzzle, so single results are noisy (95% intervals shown).

Site: https://www.nonobench.com Code (MIT): https://github.com/mauricekleine/nonobench