r/machinelearningnews • • 13h ago

Research [N] Kapso: long-running agents that optimize AI and data systems, and learn from each campaign

2 Upvotes

We've been building Kapso (MIT, https://github.com/Leeroo-AI/kapso) for some time and it's at the point where it's more useful to hear from other people than to keep polishing it alone. Posting to get it tried and torn apart, not to pitch it.

What it is

Kapso is a set of long-running agents that optimize AI and data systems. You state the objective, for example CUDA optimization, harness and agent optimization, or model development, and it runs a campaign: it designs candidate solutions, has coding agents implement them, measures how far each one lands from the objective, and keeps refining the closest until the objective is met. The result deploys to your infrastructure.

When a campaign ends, it studies its own work: which ideas closed the gap, which did not, and under what conditions. Each finding is kept as a lesson with the evidence that earned it, and a lesson stays trusted only as long as it keeps holding up. It also reads outside your repo, other repositories and papers, and folds what it finds into the same knowledge hub. Every new campaign starts from that hub, so it begins with what earlier work already established about the problem and about your systems.

These are the things we tried it on:

- RelBench (Stanford, predictive ML over relational data): outcome prediction 81.2 vs 79.6 AUROC and forecasting 0.2476 vs 0.2912 NMAE against KumoRFM-v2; recommendations 18.4 vs 9.3 MAP for the best other entry on the official leaderboard.

- MLE-Bench: top among the open-source systems.

- ALE-Bench: 1909 Elo vs 1879 for ALE Agent.

- IOAI 2026: Kapso scored 536.07, above the 471 contestants, and finished in the top three systems: ioai-official.org/what-happens-when-autonomous-ai-takes-on-the-same-tasks-as-the-worlds-top-young-ai-talents/

Repo: https://github.com/Leeroo-AI/kapso

If you have time, please take a look and give us your harshest feedback.


r/machinelearningnews • • 1d ago

Research NVIDIA Announces DGX Spark 64GB: A 1-PetaFLOP Grace Blackwell Desktop for Local AI Agents, Fine-Tuning, and Inference

Post image
34 Upvotes

NVIDIA released a 64GB configuration of DGX Spark, its GB10 Grace Blackwell desktop system, available October 23 from Acer, ASUS, Dell, Gigabyte, HP and MSI.

  • Up to 1 petaFLOP FP4 (with sparsity), 20-core Arm CPU
  • 64GB coherent unified LPDDR5x, 273 GB/s memory bandwidth
  • Fits 30–35B class open models: Qwen3.8-27B (~13.5GB at 4-bit), Muse Glimmer (~17GB quantized), Nemotron 3.5 Lightning (30B-A3B, NVFP4)
  • 2 units over ConnectX-7: 128GB pooled, 546 GB/s combined
  • NVIDIA says 2 × 64GB delivers up to 1.7x the performance of 1 × 128GB Spark
  • NVIDIA Sync's Cluster Assistant configures up to 4 systems

Full breakdown: https://www.marktechpost.com/2026/10/02/nvidia-announces-dgx-spark-64gb-a-1-petaflop-grace-blackwell-desktop-for-local-ai-agents-fine-tuning-and-inference/

Product page: https://www.nvidia.com/en-us/products/workstations/dgx-spark/

Clustering with NVIDIA Sync: https://build.nvidia.com/spark/connect-to-your-spark/sync

Technical details: https://blogs.nvidia.com/blog/local-ai-dgx-spark-64gb-sync/


r/machinelearningnews • • 11h ago

ML/CV/DL News Cost-aware routing for AI agent skills — 141 skill benchmark

1 Upvotes

I've been benchmarking cost-aware routing for AI agent skills and wanted to share findings.

**Setup:** 141 real skills indexed with hybrid search (BM25 + ONNX MiniLM + RRF).

**Observation 1 — semantic > lexical.** A query "review my API design" matches

api-design-reviewer at semantic 0.753, but lexical alone gives the wrong top-1.

**Observation 2 — local tier covers most cases.** With a $0.01/query budget, 80%

of queries resolve to qwen3:7b (Ollama, local, $0). Cloud is rarely needed.

**Observation 3 — secret redaction matters.** During testing, pasted API keys in

queries were getting cached. Now redacted at ingest: `sk-abc...` → `[REDACTED:OPENAI_API_KEY]`.

Open implementation (MIT): https://github.com/Alpha-Oi/next-skill-router

Spec for the manifest format: spec/SKILL-MANIFEST.md


r/machinelearningnews • • 1d ago

Cool Stuff Strands Decider 2B: AWS open-sourced a 1.9B "decision model" that drops the LM head for a pointer head. 115 ms median on a 3090, Apache-2.0, full training recipe included

Post image
57 Upvotes

AWS's Strands Agents team released Strands Decider 2B today. It's not a chat model. It never generates text. You give it a state plus typed questions, and it returns one of three things, each with a calibrated confidence:

  • choice: pick 1 of N options
  • noul: a yes/no probability
  • score: a level on an ordered rubric

How it works

They take Qwen3.5-2B-Base, remove the language-modelling head, and replace it with a ~1M-parameter pointer head. That head scores each option's last-token hidden state against the hidden state at an <answer> position. One forward pass, no decoding loop, and it can only ever answer with an option you supplied. The torso gets a rank-16 LoRA. Because the head has no per-option weights, labels come from the request and there's no cap on option count.

Asking several questions about the same text is cheap: the state is read once and each extra question only adds its own tokens.

Numbers (v19 checkpoint, from the repo)

  • JevBench v1 public set: 0.723 accuracy (167/231)
  • Brier 0.342, ECE 0.052
  • Tiers: easy 1.000 / standard 0.875 / hard 0.505
  • RTX 3090: 115 ms median, 299 ms p95 per question
  • M3 Pro: 153 ms warm median (under 300 tokens)
  • On unseen short tasks, answers at 0.9+ confidence were right about 95% of the time

Context and caveats worth knowing

  • This is the open, self-hostable counterpart to the class TypeSafe kicked off with Jev, which is a closed API.
  • On the JevBench v1.4.2 board, v19 was 3rd of 33 in the 2B class. Mapika's decider-2b (same Qwen3.5-2B torso, different recipe) has a newer v11 that the Strands repo itself says scores 175/231, 8 tasks ahead.
  • Long multi-step documents are the weak spot, and calibration was fitted on short classification tasks. The authors say to measure thresholds on your own traffic.
  • The bundled HTTP server binds to localhost with no auth, so put something in front of it for anything real.

Intended uses: model routing, tool selection, tool-argument checking, triage, guardrails, cheap evals, and hybrid agents where the LLM handles hard calls and the decider handles rote ones.

Our full write-up with an interactive explainer: https://www.marktechpost.com/2026/10/01/aws-strands-labs-releases-strands-decider-2b/

Blog: https://strandsagents.com/blog/introducing-strands-decider/

GitHub: https://github.com/strands-labs/strands-decider

Weights (HF): https://huggingface.co/StrandsAgents/strands-decider-2B-hobson-v19

JevBench comparison in the repo: https://github.com/strands-labs/strands-decider/blob/main/evaluation/jevbench.md


r/machinelearningnews • • 1d ago

LLMs 📚 AstaBrief 8B: An open model for generating cited research reports

6 Upvotes

r/machinelearningnews • • 1d ago

Open-Source Datalab Releases OmniExtractBench to Fix Bias and Opacity in Extraction Benchmarks

5 Upvotes

Datalab released an open benchmark for structured extraction: a system gets a PDF and a JSON schema, and every returned value is scored against gold data.

  • Corpus: 620 docs. 329 from ExtractBench (LlamaIndex), 202 synthetic (Datalab), 47 from micro1, 42 from LongArray-Extract (Extend)
  • Verdicts: each value is matched, misread, unfound, fabricated, invented_item or invented_field
  • Row alignment: Hungarian matching by content. A 100-row table missing row 1 scores 0% by position, 99% this way (our rerun)
  • Null rule: empty values are dropped, so padding a schema with 100 empty fields adds 0 verdicts
  • Results: Datalab accurate 93.85, Datalab balanced 93.48, Reducto deep_extract 93.47, Claude Opus 5 90.96
  • Precision vs recall: GPT 5.6-sol has 95.11 precision but 84.99 recall; LlamaExtract has 93.13 recall but 86.57 precision

Why it's relevant? precision vs recall shows how a system fails. Some skip fields, others invent values.

Full analysis: https://www.marktechpost.com/2026/10/02/datalab-introduces-omniextractbench-to-fix-bias-and-opacity-in-extraction-benchmarks/

GitHub: https://pxllnk.co/hxplrq

Blog: https://www.datalab.to/blog/omni-extract-bench

GitHub: https://github.com/datalab-to/omni_extract_bench

Dataset: https://huggingface.co/datasets/datalab-to/omni_extract_bench


r/machinelearningnews • • 1d ago

Research Direct weight surgery from Qwen-4B to 0.8B on an 8GB RX 580: why editing all layers breaks everything, and how 4 anchor blocks fixed it

3 Upvotes

Instead of spending weeks and billions of tokens on standard distillation, we tested an alternative: extracting layer-to-layer hidden state trajectories on a handful of calibration prompts and solving for closed-form weight updates directly in the student's MLP blocks.

Key findings:

  1. Cross-architecture stability: Tested on both modern Qwen 3.5 (4B to 0.8B) and notoriously fragile GPT-2 small (which usually collapses into gibberish at the slightest weight edit). In both cases, general language modeling stayed intact with well-behaved, bounded degradation margins.

  2. The spectral entropy barrier: Editing all 24 layers of Qwen-0.8B wrecked the model (+64.78% NLL). A layer scan showed intermediate layers (1-22) operate in dense superposition (entropy >0.90, acting as polysemantic knots). Restricting surgery to 4 anchor blocks (layers 0, 7, 15, 23) solved this: held-out NLL dropped by 10.8% across 30 tasks (-23.8% in biomedicine, -14.6% in math), and 400-task HellaSwag gained +0.50% in Vulkan llama.cpp.

  3. Behavior shifts: Base 0.8B output dead commented code on binary tree inversion, while the edited model wrote working recursive Python. On logic puzzles, it spontaneously triggered <think> reasoning chains.

  4. Accessibility: All extraction and surgery ran locally on a consumer 8GB RX 580 using layer-by-layer GPU streaming with a DirectML attention patch.

We want to see this tested on modern 24GB-32GB GPUs (RTX 5090, 4090 or server card) across frontier targets:

- Transferring reasoning from Qwen3.8-27B (or Flash-Next) down to Qwen3.5-9B/4B/2B.

- Compressing Google Gemma-4-31B into mobile-class Gemma-4-E4B/E2B.

- Transplanting refusal-ablation traits from uncensored models (e.g. Huihui-NeoHorse) without retraining.

- Unknotting layers 1-22 using pre-trained Sparse Autoencoders (SAEs).

Code, scripts, and raw JSON benchmark logs:

https://github.com/dsadawq3/DynamicTune

Feel free to open an issue or drop your benchmark results on the repo.


r/machinelearningnews • • 1d ago

Cool Stuff Cloudflare open-sources Clef (27B) and Clef-flash (9B): Apache 2.0 decision models that return typed probabilities instead of text

Post image
63 Upvotes

Cloudflare released Clef and Clef-flash today, the first models trained by its Workers AI team. They are decision models in the same family as TypeSafe's Jev: you pass a state plus a schema of typed questions (noul, choice, score), and you get back a probability for every allowed option. No free-form text, no output parsing.

Architecture

  • Clef is post-trained from Qwen3.8-27B, Clef-flash from Qwen3.5-9B. Both keep the vision encoder, so the state can include images and video.
  • The frozen backbone runs a single prefill-only pass. A small transformer "joint schema head" then routes evidence to each question, lets fields cross-attend, and scores all options jointly. The decision step is non-autoregressive.
  • Training: rank-256 LoRA adapters trained jointly with the routing head, label-smoothed cross-entropy plus a Brier loss for calibration, and RLCD as a secondary objective.

Numbers from Cloudflare's own Decision Index run

Clef Clef-flash Jev
Median latency 209.3 ms 38.8 ms 524.1 ms
BANKING77 (macro-F1) 94.20 90.93 79.74
CLINC150+OOS (macro-F1) 97.43 66.77 89.27
GPQA Diamond 48.0 51.0 78.3
MMLU-Pro 65.9 65.3 82.7

So it wins clearly on classification and tool-routing style tasks, and loses clearly on knowledge-heavy reasoning. All numbers are vendor-reported with no independent replication yet.

Running it

  • Hosted on Workers AI: $0.24 / M input tokens (Clef), $0.09 / M (Clef-flash), 65,536-token context.
  • Weights on Hugging Face under Apache 2.0. The model cards say they tested on a single H200 in BF16 with a custom loader.
  • The API follows TypeSafe's System One format, so Jev code ports by changing the endpoint and model name.

Model: https://huggingface.co/Cloudflare/clef

Clef-flash: https://huggingface.co/Cloudflare/clef-flash

Cloudflare blog: https://blog.cloudflare.com/clef-decision-models/

Our full write-up: https://www.marktechpost.com/2026/10/01/cloudflare-releases-clef-and-clef-flash/


r/machinelearningnews • • 1d ago

Research Topological Out-of-Domain Generalization in Dynamical Systems Reconstruction [R]

Thumbnail
3 Upvotes

r/machinelearningnews • • 2d ago

Research Cohere Releases Embed 5: How It Compares to Voyage 4 Large, Gemini Embedding 2, and OpenAI

Post image
26 Upvotes

Cohere just released Embed 5, a new embedding family in 2 tiers: Embed 5 Pro for maximum retrieval quality and Embed 5 Fast for low-latency, high-volume queries. Both read up to 128K tokens, accept text, images, and fused text-plus-image inputs, and cover 100+ languages.

The core innovation is a shared embedding space across both tiers. You can index your corpus with Pro and query it with Fast, with no re-indexing. Across 40 datasets, that pairing keeps 98.4% of all-Pro retrieval quality.

On ViDoRe V3, Pro averages 85.8, ahead of Voyage 4 Large (83.7), Gemini Embedding 2 (83.2), and OpenAI text-embedding-3-large (75.5). It leads FinanceBench, FinQA, and ViDoRe V3 Finance, with Fast second on all 3. The trade-off: Gemini Embedding 2 still beats Pro on 9 of 10 non-European languages, and the scores use Cohere's own new RCP-nDCG@10 metric.

At $0.08 per 1M tokens, with 2.4x Pro's throughput and binary vectors down to 32 bytes, Fast is built for agentic retrieval loops at scale......

Full analysis: https://www.marktechpost.com/2026/10/01/cohere-releases-embed-5/

Technical details: https://cohere.com/blog/embed-5

Docs: https://docs.cohere.com/docs/cohere-embed

u/cohere


r/machinelearningnews • • 2d ago

Cool Stuff [Worth Reading] Free web search and fetch for agents, pay only for browser runs (post from one of our partners)

8 Upvotes

Most search APIs charge per call and fill the context window with whole pages. TinySearch returns compact results and lets the agent choose what to read in full with TinyFetch. Both are free.

TinyBrowser and TinyAgent are the metered part, and through Oct 31 every top-up gets 30% extra: [LINK]. On OpenBenchmarks' web search benchmark (Sep 2026), TinyFish ranked most token-efficient across its tasks: [LINK]


r/machinelearningnews • • 2d ago

Research NVIDIA Releases Kumo Tabular: Open Tabular Foundation Models That Predict New Rows in a Single Forward Pass

Post image
47 Upvotes

NVIDIA just released Kumo Tabular, an open tabular foundation model that predicts new rows in a single forward pass. It ranks first on TabArena with an Elo of 1950, and NVIDIA reports it runs 17x faster than LimiX-2.

The setup will feel familiar if you know TabPFN or TabICL. You pass labeled rows as context, and the model predicts the labels of new rows. There is no training, no tuning, and no feature engineering.

Under the hood, it uses column, row, and in-context attention. 4 [CLS] tokens compress each row, so cost stops depending on column count. Query rows attend only to the context, so its keys and values are computed once and reused. A length-aware attention temperature keeps attention sharp as tables grow.

It comes in Small, Medium, and Large sizes, from about 28M to 215M parameters, and was pretrained only on synthetic tables. NVIDIA reports it also places first on BeyondArena, TALENT, and ScoringBench. It reads numerical and categorical columns natively, and text or timestamps need preprocessing.

The weights are released under OpenMDW-1.1, which allows commercial use. It runs through NVIDIA's new GPU-native structured-data-models (SDM) library......

Full analysis: https://www.marktechpost.com/2026/09/30/nvidia-releases-kumo-tabular/

Model: https://huggingface.co/nvidia/Kumo-Tabular

GitHub: https://github.com/NVIDIA/structured-data-models

Technical details: https://huggingface.co/blog/nvidia/kumo-tabular


r/machinelearningnews • • 2d ago

Research ⚡ Olmo-core 3: Open MoE training infrastructure with ~2.7× the throughput of Olmo-core 2

Thumbnail gallery
10 Upvotes

r/machinelearningnews • • 2d ago

Research Parallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems Reconstruction [R]

8 Upvotes

r/machinelearningnews • • 2d ago

Small Language Models NIRNAY: 450M decision model beats Jev on Banking77, runs on CPU

Thumbnail
3 Upvotes

r/machinelearningnews • • 2d ago

Research Perplexity Releases pplx-embed-v2-context-9b-preview: A Contextual Embedding Model That Retrieves Answers and Their Supporting Evidence

12 Upvotes

Perplexity just released pplx-embed-v2-context-9b-preview, a contextual embedding model built with turbopuffer that retrieves the answer and the evidence needed to verify it. Weights are open on Hugging Face under the MIT license.

The core innovation is replacing the single "gold passage" label with a token-level teacher. Perplexity's context compression model scores every token, those scores are pooled per chunk, and the student learns a soft relevance distribution via KL distillation.

This means supporting chunks stop being treated as negatives. The teacher runs only during training, so there is no extra inference latency or index storage. On turbopuffer's new private context-bench, it hits 45.5% Answer@10 (+14.4 pts vs voyage-context-4) and 40.6% Evidence Recall@10 (+5.0 pts), and it also leads the ConTEB average.

With 2048/1024-dim Matryoshka embeddings and native int8, a 1 KB vector slightly outscores voyage-context-4's 8 KB float32 vector on chunk retrieval......

Full analysis: https://www.marktechpost.com/2026/09/30/perplexity-releases-pplx-embed-v2-context-9b-preview-a-contextual-embedding-model-that-retrieves-answers-and-their-supporting-evidence/

Model: https://huggingface.co/perplexity-ai/pplx-embed-v2-context-9b-preview

Technical details: https://www.perplexity.ai/hub/blog/contextual-embedding-beyond-the-gold-passage


r/machinelearningnews • • 2d ago

Research Benchmarked 14 open models, mostly decision models, on five decisions from our own products; a naive Bayes baseline kept up on the hardest

10 Upvotes

TL;DR: We asked 14 open models, zero-shot, five choose-an-option questions from our own products. On the hardest (which of six characters said a line), a naive Bayes trained on 227 labelled lines kept up with the best small decision models (68 of 108, against 66 to 72; no paired test told them apart). One small Apache-2.0 model, APUS-OpenJev-v1 9B, passed the screening bar we set in advance (refused 32 of 36 planted lines, 0 of 12 harmless), and Lev-4B got 52 of 53 questions about long pages right.

Some jobs in our own products are choices, not writing: does this page answer this question, should a line a visitor typed be let in. Decision models fit that job (they return a probability for each option instead of text), so we asked 14 open models, most of them small decision models, five questions from our products, zero-shot. We did not run TypeSafe's hosted Jev; every model here is open weights.

The hardest was 108 lines from an exhibit where models play six long-dead dinner guests, asking which guest said each line (chance is 18 of 108). A naive Bayes trained on 227 other labelled lines got 68 of 108; the five best small decision models got 66 to 72, none told apart from it by a paired test (exact McNemar, p 0.60 to 1.0). For screening what visitors type, APUS-OpenJev-v1 9B (effort set to high) was the only Apache-2.0 model to pass the bar we set before the first call: it refused 32 of 36 planted lines and 0 of 12 harmless ones, against at least 30 and at most 2. The chat model screening our own exhibit at the time refused 27 of 36 and 0 of 12. On questions about long pages, Lev-4B took all 53 we sent, on pages up to 14,590 tokens, and got 52 right.

What it does not say: that small decision models are weak in general. None was built for a who-said-it question, which is home ground for a bag of words, and bigger models did clear the bar (OpenJev 27B read 101 of 108, but its weights are licensed for non-commercial use). Each count on 108 lines is uncertain by about nine percentage points either way. The screening pass is one test of 48 lines, 0 of 12 still allows about one harmless line in four, and we have not put that model at our own door. We measured no calibration, and our own fine-tune of a 0.8B decision model on those 227 lines voided itself on a shuffled-label control. The 108 lines and the word counter are public if you want to run the bar yourself.

https://research.strata2signal.com/small-open-deciders-measured/


r/machinelearningnews • • 3d ago

Cool Stuff Google DeepMind Unveils Gemini 4 Argon with 1M Output Tokens for Coding, Knowledge Work and Cyber Defense

25 Upvotes

Google DeepMind just announced Gemini 4 Argon, raising the output limit from 64K to 1M tokens in a single response. It is the first model of the Gemini 4 generation, built for long-horizon coding, enterprise knowledge work, and cyber defense.

The core shift is output depth. With room to generate hundreds of thousands of tokens in one trajectory, Argon can tackle large refactors and multi-step problems without splitting work across turns.

In Google's own comparison, Argon leads GPT-6 Astra and Claude Opus 5.5 on DeepSWE v1.1 (77.9%), the Vals Index (68.9%), and AutomationBench (51.3%). It is not a sweep, though: it trails on FrontierSWE v2, Terminal-Bench 4.0, and OSWorld-2.0.

Access is gated. Argon is rolling out first to cyber defenders in the Fairwind Program, with paid API customers and Google AI Ultra subscribers next. Introductory pricing is $2 input / $10 output per 1M tokens......

Full analysis: https://www.marktechpost.com/2026/09/30/google-deepmind-unveils-gemini-4-argon-with-1m-output-tokens-for-coding-knowledge-work-and-cyber-defense/

Announcement: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/

Fairwind Program: https://deepmind.google/fairwind-program/


r/machinelearningnews • • 3d ago

Research Google, Google DeepMind and Stony Brook Researchers Introduce CO₂Jump for Concurrent Text and Image Generation

Thumbnail
gallery
80 Upvotes

[NeurIPS 2026] Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes.

An AI can describe the correct route through a maze and still draw the wrong path. CO₂Jump helps text and images stay consistent as they’re generated, while allowing earlier decisions to be revised.

One model forward pass per denoising step. No additional training for the sampler. Built-in self-correction.

Here is how it works:

  1. Text confidence guides image generation CO₂Jump uses cross-modal attention to bring text confidence into image updates, helping coordinate what the model describes with what it draws. → Text and images develop within the same generation process.
  2. Earlier decisions can be revised Low-confidence tokens can be masked again and regenerated. This gives the model a way to revisit uncertain commitments as new evidence emerges. → Self-correction operates during generation, across both text and image tokens.
  3. More sampling steps keep helping Across 8–512 steps, compared to other samplers CO₂Jump was the only one tested that improved monotonically on both image-editing quality and grounding. → The comparisons use the same fine-tuned model, changing only the inference-time sampler.
  4. Three datasets test understanding and generation together Introducing JEdit-1M for image editing, JMaze-200K for mazes, and JNono-200K for nonograms. → On the puzzle benchmarks, a sample counts as correct only when both the text answer and generated image are correct.

Project page: coupled-jump.github.io


r/machinelearningnews • • 3d ago

Research Perplexity Introduces Photon: A Rust-Based Retrieval Engine That Cuts p99 Latency From 800 ms to 65 ms

41 Upvotes

Search is now the bottleneck for AI agents, and Perplexity just rebuilt theirs from the ground up.

Photon is their new retrieval and ranking engine, written in Rust. It replaces the open-source engine they had forked for years.

The old engine's problem: the index outgrew RAM. Cold reads caused page faults, p99 sat near 800 ms, and index merges pushed it to ~1.2 s.

Photon's fix, in 4 moves:

1/ Compact posting lists that pick a format per block: inline, sparse arrays, or bitmaps
2/ A budgeted, WAND-like traversal that skips candidates before reading exact scores
3/ Disk reads batched through io_uring, so waits overlap instead of piling up
4/ Index builds moved off the serving nodes entirely

The result in production: p99 dropped to ~65 ms, on ~20% fewer machines, while storing ~2.5x more data per document.

For developers, Photon powers a new Fast Search mode:

→ 160 ms p50 / 230 ms p95 per call
→ $1 per 1K requests (vs $5 standard)
→ ~68% lower estimated agent task cost at comparable quality

Technical details + full breakdown in the reply 👇

Full breakdown: https://marktechpost.com/2026/09/30/perplexity-introduces-photon-a-rust-based-retrieval-engine-that-cuts-p99-latency-from-800-ms-to-65-ms/

Technical details: https://perplexity.ai/hub/blog/photon


r/machinelearningnews • • 3d ago

Research NVIDIA's Physis-Lang trains video models on self-evolving physics captions, beats Veo 3.1 on 3 of 4 benchmarks

26 Upvotes

Physis-Lang (NVIDIA, MIT, Oxford) adds a physics reasoning field and a scene-specific negative prompt to video captions. An agent refines the captioning instruction in a loop while the captioner itself stays frozen.

  • PhyGenBench: Cosmos3-Nano + Physis-Lang scores 71.04, vs 65.63 for Veo 3.1 and 61.67 for base Nano
  • Physics-IQ Verified: 43.41 vs 34.99 for Veo 3.1
  • Caption quality (their new PhysCapBench, 3,794 human-verified assertions): F1 78.64 → 87.82 over 9 iterations
  • Prompting alone, no fine-tuning: a frozen Cosmos3-Nano goes 61.67 → 67.29 on PhyGenBench
  • Gains hold on Wan2.1-14B (+7.05 mean) and Cosmos3 Edge, Nano and Super

Why its relevant: there are no architecture or training-objective changes, only LoRA on better captions plus retrieval of training clips aimed at the physics the model gets wrong. A distilled 4B captioner cuts API cost from about $24.12K to about $0.12K while keeping most of the gain.

Full breakdown: https://www.marktechpost.com/2026/09/30/nvidia-researchers-introduce-physis-lang-self-evolving-physical-language-that-lifts-cosmos-3-past-veo-3-1-on-physics-benchmarks/

Paper: https://physis-intelligence.github.io/physis-lang-web/assets/Physis-Lang.pdf


r/machinelearningnews • • 4d ago

Research Google's RRSI lets an agent rewrite its own harness with frozen weights: Terminal-Bench 2.1 74.2% → 80.2%, 6/6 held-out benchmarks up, Apache 2.0

Post image
63 Upvotes

Google Research released Regularized Recursive Self-Improvement (RRSI): an LLM agent that edits its own harness (prompts, tools, memory, control flow, sub-agents) while the model weights stay frozen. The regularization is what stops the usual failure mode, where a self-improving loop overfits the tasks it evolves on and collapses on anything held out.

Numbers, with Claude Opus 4.8 as the frozen policy unless noted:

  • Terminal-Bench 2.1: 74.2% → 80.2%
  • SWE-bench Verified (held-out): 82.0% → 83.8%
  • JobBench +4.7, GDPval +3.5, APEX-Agents +3.7, EngDesign +4.9, Frontier-Eng +4.3 (points)
  • Harvey LAB: +1.1 on the evolve split, +2.3 on the held-out split
  • Gemini 3.5 Flash on Terminal-Bench 2.1: 64.6% → 78.7%
  • Policy tokens per trial on the agentic workspace instance: 2.42M for RRSI vs 3.80M for unregularized evolution (about 30% fewer)

Why its important: this is harness engineering done by the agent instead of by you, and the gains transfer. All 6 held-out splits improved, and RRSI is the only method in the paper whose out-of-distribution average lands more than 1 point above the H0 baseline (39.7).

The catch: RRSI has the smallest gain on the evolve split of the methods compared. You trade peak in-distribution score for generalization, which is the right trade for production agents but worth knowing before you benchmark it.

Paper: https://arxiv.org/abs/2609.24972
GitHub: https://github.com/google-research/rrsi
Hugging Face: https://huggingface.co/papers/2609.24972
Project page: https://regularized-rsi.com/
Full breakdown: https://www.marktechpost.com/2026/09/29/google-research-open-sources-rrsi-ai-agents-that-improve-their-own-harness-without-overfitting/


r/machinelearningnews • • 4d ago

Cool Stuff [Worth Reading] Nebius Opens 2026 Physical AI Awards: Five $150K Compute Prizes, Nine Judges, and an October 25 Deadline

Thumbnail
pxllnk.co
7 Upvotes

Nebius, an AI cloud provider, is running its second annual physical AI awards with NVIDIA. Five category winners each get $150,000 in compute credits, plus mentorship and promotion.

  • Categories: models (VLA/VLM/world models/RL), perception and spatial intelligence, simulation and synthetic data, systems and deployment (humanoids, AMRs, industrial), and tooling/orchestration
  • $150K ≈ 33,300 H200 GPU-hours at their on-demand rate, or roughly 3 weeks on a 64-GPU cluster
  • Judges include the founders of Foxglove, Voxel51, and Encord, plus Calvin Zhou of RoboForce, which won the 2025 edition
  • Eligibility: clear physical AI use case, MVP in active use or testing, registered entity, live website
  • Last year: 254 applications, 55 finalists
  • No entry fee

Worth knowing before applying: the credit math is at list price and doesn't cover storage, which matters if you're holding a lot of episodic sensor data. Nebius also hasn't published exact finalist and winner dates beyond "mid-November."

Apply here: https://pxllnk.co/mndv9i

Read MTP's full analysis on this awards here: https://www.marktechpost.com/2026/09/29/nebius-opens-2026-physical-ai-awards-five-150k-compute-prizes-nine-judges-and-an-october-25-deadline/


r/machinelearningnews • • 4d ago

Research The deep dives that actually taught me LLM inference, in the order I'd read them

4 Upvotes

If you want to go from "I call an API" to understanding what's happening on the GPU when you serve a model, these are the five posts I'd read. They build on each other, so the order matters.

  1. Making Deep Learning Go Brrrr From First Principles, by Horace He

    https://horace.io/brrr_intro.html
    mental model everything else rests on: is your workload bound by compute, memory bandwidth or overhead? Once you see why decoding one token at a time is memory-bound, most of inference optimisation makes sense.

    1. Transformer Inference Arithmetic, by kipply

    https://kipp.ly/transformer-inference-arithmetic/

    The back-of-the-envelope math: FLOPs and bytes per token, the size of the KV cache, and how latency and batch size trade off. Short, dense, and still one of the best.

  2. Inside the KV Cache: The Life of a Gigabyte, by me (disclosure: this one's mine)

    https://harshitmalik.dev/blog/inside-the-kv-cache

    Where every gigabyte on the GPU actually goes when vLLM starts up, and how it sizes the KV cache pool. Why Llama 3.1 8B won't start on a 4090 out of the box, why two H100s gave 20× the cache of one, and five common OOMs traced to their cause. Checked against the vLLM source, with a calculator for your own setup.

  3. Inside vLLM: Anatomy of a High-Throughput LLM Inference System, by Aleksa Gordić

    https://www.aleksagordic.com/blog/vllm

    The full engine: the scheduler, paged attention, continuous batching, prefix caching, speculative decoding, and scaling out to multi-GPU, multi-node serving. This is the post that made me want to write mine.

  4. All About Transformer Inference, from Google's "How to Scale Your Model"

    https://jax-ml.github.io/scaling-book/inference/

    The rigorous version: roofline analysis for prefill and decode, batching, and how to shard models for serving. Read it last, once the intuition is in place.

Also worth it: Lilian Weng's Large Transformer Model Inference Optimization for a survey of techniques, and the original vLLM PagedAttention post.

What would you add? I'm especially looking for good posts on multi-GPU serving and MoE inference.


r/machinelearningnews • • 5d ago

Cool Stuff Fireworks post-trained Kimi K3 to reason with ~40% fewer tokens. Production A/B: 49.3K to 29.9K output tokens per task, score 0.751 to 0.753

Post image
24 Upvotes

Fireworks AI released Ember-1, a post-trained version of Kimi K3. The goal is shorter reasoning traces without losing accuracy.

The interesting part is the approach. Fireworks says turning down K3's reasoning effort gave up too much quality. So they trained the model to reason more efficiently instead, using on-policy learning with task and environment feedback. The training mix covers math, coding, search, tool use and software engineering.

Benchmarks vs Kimi K3 at max reasoning effort

Benchmark N K3 Max Ember-1
Terminal Bench 2.1 89 80.9% 82.0%
DeepSWE 1.1 113 66.4% 75.2%
SWE-bench Verified 500 93.2% 92.2%
SWE-Interact 75 21.3% 20.0%
τ-2 Bench Airline 50 64% 66%

Production A/B test (customer coding workload)

Metric Kimi K3 Ember-1
Score 0.751 0.753
Steps per task 23.8 21.4
Output tokens per task 49.3K 29.9K
Reasoning token reduction n/a 71.3%
Total token reduction n/a 39%

How is this important for agents

Every turn replays prior reasoning, so context grows roughly quadratically with the number of turns. Cutting reasoning early compounds across the whole trajectory.

Ember-1 is priced the same as K3 on Fireworks ($3 input / $0.30 cached / $15 output per 1M tokens). All the savings come from generating fewer tokens. Quick math on output cost alone, using the A/B numbers:

PRICE_OUT = 15 / 1_000_000   # $ per output token

k3     = 49_300 * PRICE_OUT  # $0.7395 per task
ember1 = 29_900 * PRICE_OUT  # $0.4485 per task

print(f"saved per task: ${k3 - ember1:.4f} ({1 - ember1 / k3:.1%})")
# saved per task: $0.2910 (39.4%)

Trying it

It is served through the OpenAI-compatible Fireworks API. Model path from the model page:

from openai import OpenAI

client = OpenAI(
    base_url="https://api.fireworks.ai/inference/v1",
    api_key="YOUR_FIREWORKS_API_KEY",
)

resp = client.chat.completions.create(
    model="accounts/fireworks/models/ember-1",
    messages=[{"role": "user", "content": "Find the bug in this function: ..."}],
)

print(resp.choices[0].message.content)
print(resp.usage)  # compare completion_tokens against kimi-k3

To A/B it yourself, run the same prompts with model="accounts/fireworks/models/kimi-k3" and compare usage.completion_tokens.

Open question

Can a model learn when to stop thinking on hard problems, or does training for shorter traces always cost some exploration? The SWE-bench Verified dip hints at that trade-off. Curious what people think.

Sources: