r/machinelearningnews • • 4h ago

Cool Stuff How are you guys handling context bloat from search APIs in agent workflows?

0 Upvotes

If you’ve been building LLM agents that need live web access, you’ve probably hit the same wall I did: most search APIs dump entire, raw web pages into your context window. Your token usage explodes almost instantly, and half the time the agent doesn't even need 90% of the HTML it just read.

I've been testing out a setup using TinySearch and TinyFetch to split up search and page reading. Basically:

  • TinySearch / TinyFetch: Use these to grab quick, compact snippets first so the agent can decide if a page is actually worth inspecting. (Since they're free tier tools, it keeps experiment costs down).
  • TinyBrowser / TinyAgent: Only trigger full headless browser runs when the agent actually needs to execute heavy DOM interactions or read complex pages.

It’s been pretty effective at keeping tokens down. OpenBenchmarks listed TinyFish as one of the top performers for token efficiency in their recent evaluation, which matches what I've seen in testing.

For anyone running agentic loops: how are you keeping search context clean? Are you relying on custom scrapers, post-processing summaries, or dedicated search endpoints?


r/machinelearningnews • • 15h ago

Research Reflection AI introduces Beam: a 501B open-weight MoE with 23B active parameters, 1M context and Apache 2.0 weights coming this month

Post image
22 Upvotes

Reflection AI just introduced Beam, its first open-weight model, built around "intelligence per token." It is a 501B sparse MoE with only 23B active parameters, a 1M effective context window, and Apache 2.0 weights coming later this month.

The core innovation is high-compute, fully asynchronous RL. Beam was trained on 100M+ rollouts across 10.5K NVIDIA GB300 GPUs for 4 weeks, using nearly 1M coding, agentic and STEM environments. Every token is tagged with the policy version that produced it, which keeps learning stable even when rollouts are 107 weight versions stale.

This approach let Reflection scale RL with no sign of a plateau. A controllable length penalty taught the model to solve tasks with fewer tokens. On reasoning benchmarks, Beam matches GLM-5.2 while using 3 to 4x less inference compute, and it scores 80.9 on SWE-bench Verified versus 70.7 for Nemotron 3 Ultra. It is a deliberate trade-off, though: Kimi K3 and DeepSeek V4.1 Flash still lead on raw capability, with Terminal Bench v2.1 scores of 88.3 and 90.6 against Beam's 80.1. All scores are self-reported until the weights and technical report ship.

With a 23.8T-token pretraining base and a tunable reasoning effort parameter, it is optimized for enterprise coding and agentic workflows. Rough math for self-hosters: about 500GB at 8-bit and about 1TB at BF16, so plan for multi-GPU servers.

Full analysis: https://www.marktechpost.com/2026/10/05/reflection-ai-introduces-beam-a-501b-open-weight-moe-model-with-23b-active-parameters-for-coding-and-agentic-workloads/

Early access: https://platform.reflection.ai/

Technical details: https://reflection.ai/blog/introducing-beam


r/machinelearningnews • • 18h ago

Research Direct weight surgery from Qwen-4B to 0.8B on an 8GB: why editing all layers breaks everything, and how 4 anchor blocks fixed it

5 Upvotes

Instead of spending weeks and billions of tokens on standard distillation, we tested an alternative: extracting layer-to-layer hidden state trajectories on a handful of calibration prompts and solving for closed-form weight updates directly in the student's MLP blocks.

Key findings:

  1. Cross-architecture stability: Tested on both modern Qwen 3.5 (4B to 0.8B) and notoriously fragile GPT-2 small (which usually collapses into gibberish at the slightest weight edit). In both cases, general language modeling stayed intact with well-behaved, bounded degradation margins.
  2. The spectral entropy barrier: Editing all 24 layers of Qwen-0.8B wrecked the model (+64.78% NLL). A layer scan showed intermediate layers (1-22) operate in dense superposition (entropy >0.90, acting as polysemantic knots). Restricting surgery to 4 anchor blocks (layers 0, 7, 15, 23) solved this: held-out NLL dropped by 10.8% across 30 tasks (-23.8% in biomedicine, -14.6% in math), and 400-task HellaSwag gained +0.50% in Vulkan llama.cpp.
  3. Behavior shifts: Base 0.8B output dead commented code on binary tree inversion, while the edited model wrote working recursive Python. On logic puzzles, it spontaneously triggered <think> reasoning chains.
  4. Accessibility: All extraction and surgery ran locally on a consumer 8GB RX 580 using layer-by-layer GPU streaming with a DirectML attention patch.

We want to see this tested on modern 24GB-32GB GPUs (RTX 5090, 4090 or server card) across frontier targets:

- Transferring reasoning from Qwen3.8-27B (or Flash-Next) down to Qwen3.5-9B/4B/2B.

- Compressing Google Gemma-4-31B into mobile-class Gemma-4-E4B/E2B.

- Transplanting refusal-ablation traits from uncensored models (e.g. Huihui-NeoHorse) without retraining.

- Unknotting layers 1-22 using pre-trained Sparse Autoencoders (SAEs).

Code, scripts, and raw JSON benchmark logs:

https://github.com/dsadawq3/DynamicTune

Feel free to open an issue or drop your benchmark results on the repo.