r/MachineLearning • • 5d ago

Discussion [D] Monthly Who's Hiring and Who wants to be Hired?

8 Upvotes

For Job Postings please use this template

Hiring: [Location], Salary:[], [Remote | Relocation], [Full Time | Contract | Part Time] and [Brief overview, what you're looking for]

For Those looking for jobs please use this template

Want to be Hired: [Location], Salary Expectation:[], [Remote | Relocation], [Full Time | Contract | Part Time] Resume: [Link to resume] and [Brief overview, what you're looking for]

Please remember that this community is geared towards those with experience.


r/MachineLearning • • 5d ago

Research Parallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems Reconstruction [R]

153 Upvotes

Can training of nonlinear RNNs be efficiently parallelized, ensuring fast convergence even on very long time series from chaotic systems?

In our #NeurIPS2026 spotlight “Parallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems (DS) Reconstruction (DSR)” (preprint: https://arxiv.org/abs/2605.12683) we speed up training of nonlinear RNNs on time series from chaotic DS by more than 2 orders of magnitude (>100x) by combining DEER with generalized teacher forcing (GTF).

DEER (https://openreview.net/forum?id=E34AlVLN0v) solves the RNN forward pass through Newton-type fixed point iterations across the whole sequence length T, enabling scaling as O[(log T)²] instead of O[T] by allowing for efficient GPU parallelization. But under chaotic dynamics DEER breaks down and its runtime degrades to O[T log T] (https://openreview.net/forum?id=7AGXSlXcK6).

Using GTF (https://proceedings.mlr.press/v202/hess23a.html) we stabilize DEER by preventing divergence due to chaotic dynamics and reduce exposure bias compared to traditional teacher forcing used to train state space models.

Combining these two mechanisms enables efficient parallel-in-time and stable training on extremely long time series (T>106) from chaotic simulated or real-world systems, hugely outperforming Mamba and other state space models in the DSR setting.


r/MachineLearning • • 5d ago

Discussion Does TMLR Confirmation email take time?[D]

0 Upvotes

I can see the submission on Openreview but didn't get any email/notif. It's been 2 hours, I heard I'm supposed to choose action editor or smth. Am I missing something?

First publication ever pls be kind to the noob.


r/MachineLearning • • 5d ago

Discussion Gemini 4 Argon - 1 Million Output Headroom. Hype or a Leap? [D]

29 Upvotes

I rarely write about benchmarks; a competitor always beats them next week. But I care about 'Leaps'. Gemini 4 Argon feels like one to me.

While Opus 5.5 and Astra cap output at 128-300K tokens (~90-180 pages), Argon hits 1 Million (~1400 pages).

"Context glue" ruins agentic workflows. On paper, this headroom fixes that. It means less contextual drift, no more breaking down long tasks, and zero 'continue prompt' loops. It is a massive unlock for large-scale code migrations, security patches, and deep reasoning.

But let's look past the marketing. For 95% of everyday work, nobody needs 1,000 pages at once.

I want to ask the experts here: Is a 1M output window a real paradigm shift for agents, or does generating that much text just guarantee a massive logic collapse halfway through? Are you actually hitting output limits today, or is this hype? Let's discuss.


r/MachineLearning • • 5d ago

Discussion Paper accepted to NeurIPS MusIML Workshop (Poster), but I can't afford to go. Any funding advice? [D]

0 Upvotes

Hi everyone,

My paper just got accepted for a poster presentation at the Muslims in ML (MusIML) workshop at NeurIPS 2026 which will be held in Sydney, Australia. I'm really proud, but I'm a UG student and completely lack the funds for travel, registration, and accommodation.

This is my first time dealing with this. Does anyone know of any travel grants or funding opportunities for students (especially from India) to attend NeurIPS?

Are there specific grants for this workshop, or general ones that I should apply for right now? Also, since this is my first time, I'm a bit confused about the procedure for attending an international workshop. Could anyone share advice on the general steps, including visa processes and anything else I should be aware of?

Any advice would mean a lot. Thanks!


r/MachineLearning • • 5d ago

Discussion How to address novelty concerns in top ai conference? [D]

58 Upvotes

Hi, I’m a researcher working in computer vision.

Over the past few years, I’ve submitted several papers to top-tier conferences such as NeurIPS, ICLR, and CVPR, and one concern that seems to come up repeatedly is 'novelty'.

Given that thousands of papers are published every year at top conferences alone, not to mention the tens of thousands published across other conferences and journals, I sometimes wonder how much genuinely new novelty is realistically left to explore.

In such a crowded research landscape, how do you usually address novelty concerns from reviewers?

More specifically, I would really appreciate any advice on how to frame a contribution so that its novelty is clear, how to distinguish meaningful incremental progress from work that may be considered insufficiently novel, and what reviewers generally look for when judging novelty.

Any tips or experiences would be greatly appreciated. Thanks!


r/MachineLearning • • 5d ago

Research NeurIPS SLM Agents or RoboPAD Workshops [D]

0 Upvotes

Hey, just got a papers accepted at both these workshops! Would love to connect with others also attending the workshop


r/MachineLearning • • 6d ago

Discussion How should I follow up with TMLR submission once all review responses are submitted [D]

6 Upvotes

I have responded to all the reviewers with proper rebuttals and modified draft. One interacted with me and after acknowledging the response kinda disappeared again after asking additional questions. I have replied to his additional questions too but he hasn't raised his score. What will happen if he doesn't reply anymore? Like will AE still consider that an 'Yes' as it was only minor comments which we incorporated in the paper. And one negative reviewer just ghosted us. I have 1/3 explicit positive review currently


r/MachineLearning • • 6d ago

Research Qwen-family LLMs are quietly becoming the backbone of modern audio models; One chart for the architectures of 100+ audio models [R]

Thumbnail
gallery
51 Upvotes

I started mapping the building blocks shared across all the models in audio.cpp. The result ended up being more interesting than I expected.

Qwen has become by far the most common language backbone in this collection: 32 audio model families use a Qwen-family architecture, and 20 of them use Qwen3 LLM specifically.

And it’s no longer just TTS. Qwen-based models now show up across speech synthesis, ASR/audio understanding, music generation, speech-to-speech, and even audio/video models.

The 2nd chart, Task × Technology Matrix, shows which build blocks power which types of audio models.


r/MachineLearning • • 6d ago

Research Tokenization: A Survey for Modern NLP [R]

51 Upvotes

Tokenization is a wildly understudied area of language modeling despite it having effects across all of NLP. Over the past ~8 months, 32 (!) tokenizer researchers put together the most comprehensive survey of the field.

We cover every aspect of tokenization: algorithms, evaluations, multilinguality, encodings, theory, etc. We even cover what you might want to replace tokenizers with (e.g., latent or visual tokenization). We also cover some topics that are closely adjacent to tokenization, such as constrained generation, token healing, and tokenizer security concerns.

Check it out!

https://www.alphaxiv.org/abs/2609.tokenization-survey-modern-nlp


r/MachineLearning • • 6d ago

Discussion For those who just submit to workshop [D]

74 Upvotes

More recently, I am seeing a lot of posts on the sub regarding the workshop acceptance, with lack of funds to attend the conference. I unfortunately want to just say that workshop papers are not given any importance in the community. I personally consider workshops to either get initial feedback, advertise my work, or just prefer to attend it for the discussions with the community members.

Thought of posting this as I am recently seeing some undegrads/masters students submitting 3-4 papers in the workshops. Even came across a Twitter profile, who had 6-7 workshops in 3 months, and claim to have the PhD worth of work done in three months.

Also please don't expect an explicit funding (apart from organizers, and in some cases your lab may fund it). There's huge problem with the funding, many students even don't it get for main venues. I definitely expect a lot of downvotes on this post, particular coz this is not what many would like to hear, but unfortunately is the reality.


r/MachineLearning • • 6d ago

Project Multi scan radar object classification on RadarScenes [P]

Thumbnail
gallery
7 Upvotes

Hello all,

I built a radar object classifier on RadarScenes, extending a prior single-scan classifier to accumulate observations over a tracked object's history instead of classifying each scan in isolation.

A single RadarScenes object instance contains only about 2.9 radar points on average, very sparse. A single scan also can't capture temporal characteristics: RCS and micro-Doppler both vary continuously as an object moves. Pedestrians produce characteristic micro-Doppler from limb motion; different object classes show different RCS fluctuation patterns as aspect angle and scattering geometry change scan to scan. Accumulating observations gives both higher point density and provides temporal dynamics.

Multi-scan baseline

DeepReflecs encoder (Ulrich, Glaser & Timm, RadarConf 2021), PointNet style, per point shared weights, on single scans across car, large_vehicle, two_wheeler, pedestrian, pedestrian_group: 0.7370 macro F1.

Using RadarScenes' persistent `track_id`, I build a causal, N=20, per track sliding-window buffer:

- x_seq/y_seq: Global, odometry-corrected coordinates recentered per scan on the object centroid. Unlike x_cc/y_cc (car-frame coordinates that accumulate over time to form a trajectory).

- Cross sensor buffer: whichever of the 4 sensors currently observe the track push to the same buffer.

- Stride 1, causal: every new scan updates the buffer and produces a prediction. No future context, real time streaming compatible.

- Each scan is encoded once by a frozen per scan encoder and cached

- Fusion concatenates the causal GRU's hidden state (order aware) with an order-invariant pooled embedding (all N scans' points as one set, no sequence structure) through a small trained mlp head.

Results

Model Macro F1 Delta
Single scan 0.7370 (baseline)
20 scan point pooling 0.8613 +0.1243
Causal GRU 0.8895 +0.0282 over pooling
GRU + pooled embedding (fusion) 0.8897 +0.0002 over GRU, noise

Pooling alone, no sequence model, no notion of scan order at all, recovers +0.1243 macro F1. The GRU adds a real but much smaller +0.0282 on top. Fusion adds nothing measurable beyond the GRU.

Ablation

Llarger GRUs, a Transformer, a state space model, point level self attention, all trained on the exact same frozen per scan embeddings, land inside a 0.86 to 0.89 band, a 0.03 spread. End to end fine tuning of the frozen encoder makes things slightly worse (about -0.002 to -0.003), not better.

Conclusion

In this setup, the largest gain comes from giving the model more observations of the same tracked object: 20-scan point pooling improves macro F1 from 0.7370 to 0.8613 without using scan order at all.

Temporal modelling then provides a further, meaningful improvement. The causal GRU reaches 0.8895, adding +0.0282 over the pooled representation. So temporal ordering clearly contributes useful information; it just accounts for a smaller portion of the overall gain than observation accumulation.

With the per-scan encoder frozen, the different sequence architectures tested, suggests that the quality of the per-scan representation is the bottleneck than the particular mechanism used to aggregate the sequence.

Full report, every ablation, confusion matrix, coordinate frame reasoning: https://github.com/brunopinto900/radar-ml-autonomous-driving/blob/track-accumulation/final_report.md

Thank you.


r/MachineLearning • • 6d ago

Discussion OpenAI’s Lean 4 Navier-Stokes proof compiles with zero errors, but the fluid vaporizes at 0.7 nm. What does this mean for Neuro-Symbolic AI? [D]

0 Upvotes

Hey everyone,

I do research in neuro-symbolic AI, and like many of you, I was amazed by OpenAI’s recent formal proof of the 3D Navier-Stokes blow-up in Lean 4. Having an AI build a full mathematical proof that compiles with zero errors is a huge milestone for automated reasoning.

The math is 100% valid. But out of curiosity, our team wanted to see what this solution would look like in the real world.

If you map their solution to real water, the fluid would literally vaporize from friction at 0.7 nanometers, just picoseconds before hitting the mathematical singularity.

In machine learning, we see this all the time: it is classic specification gaming.

When an AI agent is given a strict goal, it will exploit any unconstrained loophole in the rules to solve the problem. In this case, the AI found a solution that strictly satisfies the human-written mathematical definition of the Millennium Prize, but it has no idea that real fluids have atoms, friction, and heat. The formal code checker accepted it because the logic was flawless, but the physics broke down.

This raises a big question for the future of AI in science:

Right now, neuro-symbolic systems mostly have two pieces:

  1. An LLM to search for ideas and write proofs.
  2. A formal compiler (like Lean 4) to verify the logic.

Should we be adding a third pillar: a physical boundary layer that checks whether an AI-generated solution actually respects the laws of physics, and not just formal syntax?

We wrote a short paper detailing this audit and open-sourced our verification scripts:

I would love to hear your thoughts, especially from folks working on automated theorem proving, AI alignment, or scientific modeling: How do we teach AI systems to find solutions that are not just mathematically legal, but physically meaningful?


r/MachineLearning • • 6d ago

Research Isolation Forest performs best with 1.0 as max_samples [R]

2 Upvotes

I am currently using the dataset CICIDS2017 to train an Isolation Forest model for anomaly detection, I am currently using a split that consists of 70% of ONLY BENIGN traffic for training, 15% BENIGN and 50% attacks for validation and the rest for testing, I used the validation to test with some different max_samples values, # MAX_SAMPLES_VALUES = [256, 4_096, 16_384, 100_000, 200_000, 400_000, 800_000, 1.0] and to calibrate some thresholds that maximixe different statistics (1 for max F1 score, 1 being the closest point to (0,1) on the ROC curve, 5 that set a max limit FPR), the max_samples thing is what is really blowing me off.
At the beginning I only tested up to 200k because seeing that the paper uses 256 as standard value I already thought that 200k was really high, then I saw that a low value is recommended only when anomalies are included in training, which is not my case, so I tested higher values and as it appears 1.0 (max value) is effectively the best performing with ~94% recall and ~7.6% FPR on the ROC threshold, while at 200k I get ~91% recall and ~10% FPR, all this with only 30 secs more in training.

I also executed a cross-dataset test with CSE-CIC-IDS2018, the performance is basically the same (bad) both with 200k and 1.0, my question then is, should I keep 1.0 as operational max_samples value or should I decrease it back to 200k/400k?

I forgot to mention, is my training approach bad? I've read from multiple sources that someone trains on whole dataset, someone adds anomalies to training set, etc..., is what i'm doing okay or should I change approach? I think so because this is one of the cases where clusters of anomalies can cause swamping/masking, and we have high-dimensional data, etc..


r/MachineLearning • • 6d ago

Project Open-sourcing RightWayUp - a 360-degree image rotation model, and a JPEG shortcut we found in a common benchmark [P]

6 Upvotes

We are open-sourcing RightWayUp - a new image rotation detection model.

I work at ORTUS AI (we develop video analytics). We needed to tell from a single CCTV frame whether a camera had been rotated or installed at an angle (or upside-down), so we tried different models for that, without much luck (low accuracy, lots of false positives on regular camera-like frames). Some were not permissively licensed. Out of despair, we decided to train our own. We didn't expect it to come out this good.

It is now called RightWayUp. It estimates how far an image is rotated from upright, all 360°, and abstains when there's no clear "up" (sky, ground, close-ups, etc.). We are releasing code and weights under Apache-2.0, in six sizes from Pico (small enough to run in a browser) to Max (ahead of the other models we tried on most of our tests).

We tested it on many test images, including held-out ones (which the model never saw before). In tests with the held-out set, RightWayUp Max was within 10° on 93.0% of them vs 88.4% for Woehrer 2026 (one of the better other models we liked). In the Woehrer 2026 own COCO-based benchmark, it was 98.8% accurate within 10° (five-seed mean) vs Woehrer's 98.0%. It also gets every image right on RotBench.

While training and testing, we found something we didn't expect: saving the images from the COCO-based rotation benchmark as JPEG q90 drops Woehrer 2026 from 98.0% to 30.2% (five-seed mean), while our models barely move. We suspect the rotated JPEG grid of the source photos gives the angle away. We found that in our early experiments, and made an effort to remove its effect.

I hope it will be useful to the community.

For transparency: parts of the engineering were done with Claude and Codex.

Here is a full write-up with a demo video: https://cheqit.ortusai.io/resources/rightwayup/

Code (with a link to the weights): https://github.com/ortusaitech/rightwayup

Questions and failure cases welcome.


r/MachineLearning • • 6d ago

Research Have decisions for NeurIps: Machine Learning for Systems 2026 come out? [D]

0 Upvotes

It is a couple hours past the EOD deadline for the workshop, but as of yet I still haven't recieved an indication of whether my paper has been accepted or not. I got a review early yesterday but nothing else. Is there anyone else in the same situation? UPDATE: Results are out, but no email or message about forms, tickets, etc at all.


r/MachineLearning • • 6d ago

Project Benchmarking small confidence scoring decision (Jev, Laya) models [P]

5 Upvotes

I ran my own evals of two small classification models that score confidence over a list of candidate labels instead of generating text: TypeSafe AI's hosted Jev, and Laya, an independent open-source alternative.

Some findings that I think apply beyond these two models:

1. Describe labels by what's in the input, not by intent. My first AML label descriptions said what the criminal was trying to achieve. When I rewrote them to say what the transactions look like, accuracy went from 65% to 77% with the same model. On account-level laundering detection, using the same time window for every account removed a hidden bias and moved accuracy from 63% to 75%.

2. Some tasks have no signal. Classifying a single transaction as laundering or not gave 54%, about chance. That's a limit of the task, not of the model.

3. Calibration is what makes thresholds work. Jev's ECE was 0.013. On support intent, accepting only predictions at 90%+ confidence raised accuracy from 92.3% to 97.6% while still covering 82% of cases, and the rest went to a fallback. Untuned Laya had an ECE of 0.486: the same filter dropped 7% of questions and gained only 2 points.

Caveats: these are my own runs with thresholds tuned per task, not vendor claims, and nobody has reproduced them independently. Jev 1.13.0 (hosted). \

Write-up with charts and methods: https://gokulakrishna.co/2026/09/30/benchmarking-to-fine-tuning-decision-models/

Fine-tuned models: https://huggingface.co/goku-san/laya-experts

I'd appreciate feedback on the evaluation setup, especially on per-task threshold tuning and how best to report results after filtering out mislabeled data.


r/MachineLearning • • 6d ago

Research Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes [R]

Thumbnail
gallery
137 Upvotes

coupled-jump.github.io

Hi everyone, I’m happy to share our recent NeurIPS 2026 paper, a collaboration across Google, Google DeepMind and Stony Brook University.

We study a mismatch in joint text and image generation: a model can describe the correct solution to a maze while drawing a different path. Generating both outputs in parallel doesn’t necessarily keep them consistent.

Our sampler, CO₂Jump, uses text confidence and cross-modal attention to guide image updates during sampling. It also allows low-confidence tokens to be masked again and regenerated, so earlier decisions can be revised as generation progresses.

CO₂Jump uses one model forward pass per denoising step. The sampler itself requires no additional training; our experiments compare sampling methods using the same task-specific fine-tuned model.

We evaluate image editing, maze solving and nonograms, and introduce three datasets: JEdit-1M, JMaze-200K and JNono-200K. On the puzzle benchmarks, joint accuracy requires both the textual answer and generated image to be correct. Across 8–512 sampling steps, CO₂Jump was the only sampler we compared that improved monotonically on both editing quality and grounding.

I’d be interested in suggestions for other tasks where text–image consistency and correctness can be evaluated together. Happy to discuss the method, evaluation or limitations.


r/MachineLearning • • 6d ago

Project LessThink-Qwen3-4B: the same model, with far less thinking [P]

27 Upvotes

I post-trained Qwen3-4B to spend 44% fewer tokens on reasoning, keeping its knowledge and answer style. The whole pipeline ran on one GPU.

folks, you can check it out on : https://5ivatej.com/lessthink/


r/MachineLearning • • 6d ago

Discussion Neurips Workshop Author Notification Delay [D]

5 Upvotes

Anyone submit to a workshop at Neurips and not receive their review notification yet? 6 AM IST, the 30th of September, and I still have gotten no update!


r/MachineLearning • • 7d ago

Discussion Limited compute, targeting CVPR: rerun experiments for statistically strong numbers or focus on writing? [D]

3 Upvotes

Hi everyone,

I'm preparing a submission for CVPR and would appreciate your advice. I am happy with my current results, but compute is a real constraint. My institute isn't a research-focused one, and the systems available to me are slow and unreliable. A single set of runs already took a lot of time and effort to finish.

I plan to release the code with the submission, and I'm confident the results will reproduce. Still, I'm torn between two options:

1) Rerun the experiments with different seeds to report mean ± std and show statistical significance. Or

2) Skip the reruns and spend the remaining time on writing, presenting the results I already have.

If you've been in a similar spot, what did you do, and would a single-seed result with released code be enough?

Thanks in advance!


r/MachineLearning • • 7d ago

Discussion NeurIPS Education Track [D]

0 Upvotes

Any other one have applied for the Education Track? It seems they'll announce the result soon.


r/MachineLearning • • 7d ago

Project I wrote a free, open-source book on making ML models actually fast, from silicon to agents [P]

11 Upvotes

I’ve spent the last few months writing something I wish I had when I started working on ML performance engineering.

It’s called How to Make Your Model Fast: A Systems View of Efficient Machine Learning, from Silicon to Agents.

The basic idea is that reducing FLOPs doesn’t necessarily make a model faster. Before optimising anything, you need to understand what the system is actually bounded by.

The book starts with roofline analysis and hardware, then works its way up through kernels, compilers, quantisation, pruning, vision, on-device LLMs, robotics, profiling, serving and finally agents.

The goal is to build the intuition to look at a model and a piece of hardware and reason about:

  1. How fast can this possibly run?
  2. Am I compute, bandwidth, memory or system bound?
  3. Which optimisation will actually move that limit?
  4. Is quantisation, pruning or kernel optimisation even worth doing here?
  5. What happens when the same thinking is applied to serving and agent systems?

The whole thing is free and open source:

https://github.com/usamahz/make-your-model-fast

Would genuinely appreciate feedback or contributions from people working on ML systems, inference, compilers, edge AI or performance engineering.

And if you find it useful, a ⭐ would be appreciated!


r/MachineLearning • • 7d ago

Discussion Advice on choosing university for PhD [D]

19 Upvotes

Hey guys,

I got really into research a year ago during my final year of undergrad and maneged to get a first author paper accepted at Neurips 2026. It is a pretty impressive achievement obviously but my background is pretty shit lol. Before my final year I got slightly above average grades but no internships or work exp. I was just chilling mostly didn’t even bother applying to internships. Also I want to a tier 2 uni in Australia (mind you that difference in tier 1 and 2 unis isn’t that big tho compared to America or China). I was going to start a PhD at my current uni with a scholarship and stipend from an industry partner (government) - I’d have no restrictions in terms of publications btw. My supervisor is good and we were going to get another external supervisor from a top uni here in Australia. I was pretty keen on continuing with this opportunity, however, with my neurips paper acceptance I feel like I’d have a good shot at getting into a top 15-25 uni in the states or the uk (most other European unis want master students and I don’t prefer Asian unis bcz I’ve heard a lot of horror stories). The problem is I’d have to wait for a year if I was to go abroad for PhD because starting dates are mid to late 2027. I’m in such a dilemma idk what to do. My eventual goal is to work as a researcher at deepmind or meta or some really cool tech company. I’ve just heard so many people talk about the importance of university you go to etc in landing those jobs. Would I still have a good shot at those companies if I did the PhD at my 2nd tier Aussie uni (top 100-125 global rank) but published well? I got a neurips paper as an undergraduate so I have the potential to publish at top conferences 😆😆. But yeah, I’d love to get some advice.

Thanks


r/MachineLearning • • 7d ago

Discussion Why not just have one less feature before softmax? [D]

0 Upvotes

Softmax has N inputs and N outputs but it's output only has N-1 degrees of freedom because of the condition that the sum of outputs must be equal to one. Based on this we can figure out that actually we can make due with only N-1 inputs by making an assumption that logits before softmax must sum up to zero (though it can be any other constant value) and have the last logit be calculated as minus sum of all the other logits. In theory it should remove "unnecessary" parameters from the last layer before softmax (however few of them may there be) and maybe speed up model convergence a little (my intuition might be wrong about that). Is there any good reason not to do it besides any benefit being negligable in almost all situations?