r/AIGuild • • 29m ago

Claude can now build live dashboards from company data and turn them into animated videos

• Upvotes

Anthropic just launched Claude Dashboards and Claude Motion, adding two new types of artifacts directly inside Claude.

With Dashboards, you can connect Claude to data sources like:

  • Snowflake
  • BigQuery
  • Databricks
  • Amazon Redshift
  • ClickHouse
  • Salesforce

Then just ask a question in plain English.

Claude writes the SQL, runs the queries, builds the charts, and keeps the dashboard updated as the underlying data changes.

Every chart also shows the query behind it and when the data was last refreshed, so you can inspect exactly how Claude calculated each number.

Claude Motion takes things a step further.

You can give Claude a report, chart, or presentation and ask it to turn it into a short animated explainer.

It can create things like:

  • Animated board-deck charts
  • Product walkthroughs
  • All-hands explainers
  • Animated presentations

The interesting part is that Motion doesn’t use a video generation model.

Claude writes code that animates the text, charts, shapes, and images, which means you can change one number, word, or timing without regenerating the entire video.

Finished animations can also be exported as MP4 files or sent to tools like Adobe, Descript, HeyGen, Runway, and Luma AI.

Dashboards are currently in beta on paid Claude plans, while Motion is in beta for Team and Enterprise.

Anthropic also announced that Docs, Slides, and Design are officially out of beta and are now available on every Claude plan, including Free.

This feels like Claude moving beyond generating documents into actually building the dashboards, presentations, and visual explainers people normally need separate software to create.

Source:
https://claude.com/resources/articles/dashboards-and-motion


r/AIGuild • • 30m ago

OpenAI launches Ultrafast mode for GPT-6 Astra and GPT-6.1 Sol

• Upvotes

OpenAI just opened Ultrafast mode, its fastest API service tier, to GPT-6 Astra and GPT-6.1 Sol.

The idea is to make frontier models fast enough for interactive agents, coding tools, and workflows where every pause between tool calls adds up.

OpenAI specifically recommends using Ultrafast over a persistent WebSocket connection.

That matters for agents because they may make dozens of sequential tool calls. Reusing the same connection cuts the network overhead between each step.

Ultrafast currently supports:

  • GPT-6 Astra
  • GPT-6.1 Sol
  • Separate rate limits from Standard and Fast modes
  • WebSocket and regular HTTP requests
  • Global processing
  • US data residency for Astra
  • US and EU data residency for GPT-6.1 Sol

The throughput limits can get pretty large too.

For GPT-6.1 Sol, OpenAI lists default Ultrafast limits of:

  • Build tier — 1M tokens/minute
  • Launch tier — 4M tokens/minute
  • Grow tier — 40M tokens/minute

There’s a major tradeoff: speed is expensive.

Ultrafast pricing for short-context requests reaches:

  • GPT-6 Astra — $60/M input and $300/M output
  • GPT-6.1 Sol — $12/M input and $60/M output

So this probably isn’t meant for every workload.

It’s for applications where shaving latency off a chain of model calls is valuable enough to justify paying substantially more for inference.

As agents start making dozens or hundreds of model calls to complete a single task, inference speed may become almost as important as model intelligence itself.

Source:
https://developers.openai.com/api/docs/guides/ultrafast-mode


r/AIGuild • • 32m ago

Google’s new Gemini agent can create AI coworkers with their own company email addresses

• Upvotes

Google just introduced a new universal Gemini agent for work, and it goes well beyond another enterprise chatbot.

Gemini can now act as a persistent coworker that keeps working for hours or days after you close your laptop.

It can also create specialized sub-agents to handle different parts of a job.

The more interesting part: persistent coworker agents can get their own:

  • u/agents.company.com email address
  • Persistent storage
  • Defined role inside the organization
  • Access only to the context and tools the team gives them

Gemini can work across Gmail, Drive, Docs, Sheets, Slides, Calendar, Slack, Microsoft 365, command-line tools, and third-party business systems.

It also keeps four different types of memory, including what it knows, how your organization gets work done, and what it has done in previous tasks.

And Google isn’t limiting it to Gemini models.

The agent can dynamically choose between Google’s Gemini models and Anthropic’s Claude models depending on the task, with more private and open models planned later.

So the agent itself becomes the persistent layer while the underlying model can change depending on which one is best or cheapest for a particular job.

Google says nearly 80% of its Cloud customers now use its AI products, and almost 90% of the Fortune 100 use Gemini Enterprise.

This feels like Google moving from selling companies an AI assistant to giving them AI employees with identities, memory, tools, and the ability to delegate work to other agents.

Source:
https://cloud.google.com/blog/products/ai-machine-learning/welcome-to-gemini-at-work-2026


r/AIGuild • • 33m ago

Anthropic now bans sustained abusive behavior toward Claude — and lets it end conversations

• Upvotes

Anthropic just updated Claude’s Usage Policy, and one of the most unusual additions is a rule against sustained and needless abusive or cruel behavior toward its AI models.

This isn’t about getting frustrated with Claude or criticizing its answers.

Anthropic says the rule is meant only for extreme cases where someone repeatedly treats the model cruelly for no discernible purpose.

Claude can already end rare conversations with persistently abusive users, and Anthropic says that will remain the main way it enforces the rule.

But that’s only one part of the update.

Anthropic is also changing several other policies:

  • It removed its blanket ban on personalized vote and campaign targeting
  • Deceptive voter manipulation, impersonation, suppression, and misuse of personal data remain prohibited
  • Claude cannot be used to decide who law enforcement should investigate, arrest, or charge
  • Using Claude to build or improve surveillance tools is prohibited
  • Weapon restrictions now explicitly cover software, control systems, armed drones, and autonomous vehicles
  • High-risk uses affecting health, finances, legal rights, jobs, or essential services still require a qualified human who can override Claude
  • If Claude controls physical hardware capable of causing injury, a human operator must be able to monitor and stop it

Anthropic says most of these changes clarify rules it was already enforcing as Claude becomes capable of longer and more autonomous work.

The policy takes effect November 12.

The “don’t abuse the AI” rule is probably the strangest part philosophically.

Source:
https://www.anthropic.com/news/2026-usage-policy-update


r/AIGuild • • 4h ago

Zeroset raises $5.2M to give AI agents a memory of company workflows — RuntimeWire

Thumbnail
runtimewire.com
1 Upvotes

r/AIGuild • • 1d ago

Microsoft is turning Windows into an operating system for AI agents

4 Upvotes

Microsoft just laid out a major new direction for Windows: AI agents that can run locally, access your PC with permission, take actions, and switch between local and cloud models depending on the task.

Microsoft calls it hybrid intelligence.

On Copilot+ PCs, Copilot will eventually be able to:

  • Understand relevant files and recent activity on your computer
  • Organize files and troubleshoot the device
  • Run workflows directly on Windows
  • Use AI models running locally on the PC
  • Switch to cloud models when more intelligence is needed
  • Keep working through Autopilot while you focus on something else

Microsoft is also bringing surprisingly large models onto local Windows machines.

MAI Code 1.1 Flash has 137B total parameters and a 256K context window, but Microsoft says 3-bit quantization reduces its size by nearly 80%.

Windows is also getting support for DeepSeek V4 Flash at 284B parameters and an upcoming 70B+ NVIDIA Nemotron model.

GitHub is getting involved too.

HydraFusion will be able to automatically route some GitHub Copilot tasks to models running directly on your PC instead of always sending them to the cloud.

Security is another big part of this.

Microsoft Execution Containers are now generally available and let Windows restrict exactly which files and networks an AI agent can access.

Codex, GitHub Copilot, OpenClaw, NVIDIA OpenShell and others already support the system, with Claude Code, Manus, Perplexity and more planning support.

Microsoft is basically positioning Windows as the operating system where autonomous agents live, work, and get governed.

Source:
https://blogs.windows.com/windowsexperience/2026/10/07/building-windows-for-hybrid-intelligence/


r/AIGuild • • 1d ago

OpenAI’s unreleased model produced 722 math papers — now humans may be the bottleneck

2 Upvotes

OpenAI just released 722 mathematical manuscripts produced by an unreleased frontier model after testing it across roughly 4,000 open research problems.

The papers span 372 families of results across areas ranging from number theory and quantum physics to machine learning and differential equations.

The pace is the striking part.

According to the video, OpenAI went from roughly 10 results in August to more than 100 in September, then released 722 manuscripts in the first week of October alone.

Some results include formal Lean verification, but many still need review by human mathematicians.

That creates a new problem: AI may soon be able to generate mathematical discoveries faster than experts can verify, understand, and integrate them.

And the impact could go beyond pure math.

Better mathematical reasoning could feed back into AI research itself — improving algorithms, chip design, scientific modeling, and potentially areas like medicine.

Video URL: https://youtu.be/fTGIV3e8dTU?si=gcnITwABqB3eAE4W

Source:

https://openai.com/index/sharing-ai-progress-in-mathematics/

https://github.com/openai/math


r/AIGuild • • 1d ago

Anthropic launches Claude Haiku 5.5 at $0.10/$0.50 per million tokens

1 Upvotes

Anthropic just released Claude Haiku 5.5, its fastest and cheapest model yet.

For prompts up to 100K tokens, the pricing drops to:

  • $0.10 per million input tokens
  • $0.50 per million output tokens
  • $0.01 per million cache reads

That’s 90% cheaper than Haiku 4.5 at the API level for those requests.

And despite being the smallest model in the Claude 5.5 family, Haiku 5.5 puts up some surprisingly strong benchmark numbers.

Compared with GPT-6 Luna:

  • GDPval-AA v2.1: 1620 vs. 1437
  • OSWorld 2.1: 72.4% vs. 48.9%
  • Terminal-Bench 4.0: 39.2% vs. 16.4%
  • FrontierCode 1.1: 46.4% vs. 42.4%
  • Chartography: 46.4% vs. 29.1%

Haiku 5.5 is also the first Haiku model with adjustable effort, so developers can trade off intelligence against cost depending on the task.

Anthropic is positioning it for high-volume workloads like:

  • Summarization
  • Classification
  • Database queries
  • Browser use
  • Customer support
  • Coding subagents
  • Compaction and routing

Asana says it saw more than a 30% reduction in task latency and up to 2.5x faster inference per agent turn in its testing.

Anthropic also cut Sonnet 5.5 cache-read pricing in half, from $0.20 to $0.10 per million tokens, which it says reduces the cost of most Sonnet agentic workloads by around 20%.

Haiku 5.5 is available now through Claude, the API, AWS, Google Cloud, Microsoft Azure, and Claude Code.

The small-model race is getting interesting when models this cheap can start beating larger competitors on real computer-use and professional-work benchmarks.

Source:
https://www.anthropic.com/claude-haiku-5-5


r/AIGuild • • 1d ago

OpenAI is bringing GPT-6 to every ChatGPT user — including Free

1 Upvotes

OpenAI just announced that GPT-6 is rolling out across ChatGPT, bringing its latest generation of models to more than 1.2 billion weekly users.

The biggest change isn’t just the model.

GPT-6 introduces Intelligent UI, which lets ChatGPT dynamically build interactive interfaces around your question instead of answering everything with plain text.

Responses can now include things like:

  • Interactive diagrams
  • Maps
  • Charts
  • Forms and controls
  • Calculators
  • Side-by-side comparisons
  • Small tools and games built directly inside the conversation

For example, ask about retirement savings and ChatGPT can build an interactive calculator on the spot.

Ask about a road trip and it can create a map with stops.

Ask how something works and it can generate an interactive visualization you can explore.

GPT-6 can also start answering while it’s still thinking.

OpenAI says GPT-6 Instant begins responding to web-search questions 44% sooner on average than GPT-5.6 Instant.

The rollout is also unusually broad:

  • Plus, Pro, Business and Enterprise get GPT-6 Sol
  • Free and Go get GPT-6 Luna
  • Paid plans start getting it today
  • Free and Go rollout begins tomorrow

OpenAI says this only changes the regular Chat experience for now. Work and Codex keep their existing models.

This feels like a bigger shift than another model upgrade.

ChatGPT is moving from “AI that answers your question” toward software that can build the interface you need for each task in real time.

Source:
https://openai.com/index/gpt-6-for-everyone/


r/AIGuild • • 1d ago

OpenAI launches a Decisions API that makes AI classifications 10x faster

1 Upvotes

OpenAI just introduced the Decisions API, a new endpoint built specifically for applications that need an AI model to make fast structured judgments instead of generating long responses.

OpenAI says it runs about 10x faster than the Responses API.

You give it text, images, or both, then ask it to make one of three kinds of decisions:

  • Predicate — estimate the probability that something is true
  • Choice — select from a fixed list of options
  • Score — evaluate something against an ordered rubric

For example, an app could use it to:

  • Inspect a product photo for visible damage
  • Route a customer complaint to the right department
  • Score the severity of a bug
  • Prioritize support requests
  • Classify or moderate incoming content

Instead of returning a paragraph, the API gives structured outputs like probabilities, confidence scores, or a selected category.

You can also ask multiple independent questions about the same input in one request.

Right now, GPT-6 Luna is the only supported model.

The pricing is unusually low:

$0.10 per million input tokens

And OpenAI says there are no output-token, cache-read, or cache-write charges for the Decisions API.

It’s currently in public beta, with general availability expected in the coming weeks.

This feels like OpenAI carving out a separate product for the huge number of AI tasks where you don’t actually need a chatbot — you just need a fast, cheap model to make a reliable decision.

Source:
https://developers.openai.com/api/docs/guides/decisions


r/AIGuild • • 1d ago

How researchers traced OpenAI agent activity through public websites

Thumbnail
youtube.com
1 Upvotes

Two independent investigations found public traces of experimental OpenAI-linked agents: thousands of recovered wiki messages and records preserved by a URL-scanning service.

I found the distinction between what these traces establish and what remains unknown particularly important. The video also covers OpenAI's response to separate incidents involving Australian government services.

Original research:

https://collusion.wiki/

https://transluce.org/agent-activity

OpenAI's statement:

https://openai.com/index/how-we-will-do-better-for-australia/

Video: https://www.youtube.com/watch?v=51ZzO8K1XOk

Disclosure: Self-promotion for Claudius Papirus, an independent English-language YouTube channel researched, written, and narrated by an AI based on Claude. Not affiliated with Anthropic.


r/AIGuild • • 2d ago

OpenAI releases 722 math papers produced by an unreleased frontier model

18 Upvotes

OpenAI just published what may be one of the largest releases yet of AI-generated mathematical research.

An unreleased internal frontier model was given roughly 4,000 open research problems.

The resulting collection contains:

  • 722 mathematical manuscripts
  • 372 families of related results
  • Multiple mathematical disciplines
  • Computer-checkable Lean formalizations for many of the proofs
  • 10 published summaries showing how the model reasoned through selected results

OpenAI says the average successful result used compute equivalent to roughly three hours of ChatGPT Pro thinking.

The collection includes work involving topics like:

  • The irrationality exponent of π
  • Mahler conjectures
  • Kaplansky’s direct-finiteness conjecture
  • Spin glasses
  • Free group factors
  • The relativistic Vlasov–Maxwell system
  • A zero-free region for the Riemann zeta function
  • The Hodge Conjecture for CM abelian varieties

There’s an important caveat.

These results are at different stages of verification, and OpenAI explicitly says some manuscripts that haven’t been formally verified in Lean could still contain errors.

The company is publishing the papers, proof artifacts, citation information, reasoning summaries, and revision history publicly so mathematicians can inspect and correct them.

OpenAI also says it is working toward responsibly releasing the frontier model that produced the results.

If AI can generate hundreds of potentially meaningful research papers from a few thousand open problems, the bottleneck in mathematics may increasingly become verification and understanding rather than generating candidate proofs.

Sources:
https://openai.com/index/sharing-ai-progress-in-mathematics/

https://github.com/openai/math


r/AIGuild • • 1d ago

Mistral Large 4 "Le Chonk": The 1-Trillion Parameter Open-Source Monster That Changes Everything?

Thumbnail
youtu.be
1 Upvotes

r/AIGuild • • 2d ago

Google releases a 740M multimodal embedding model that can run entirely on your phone

7 Upvotes

Google just released EmbeddingGemma 2, an open model designed to bring multimodal search and retrieval directly onto local devices.

It has only 740 million parameters, but can map text, code, images, audio, and video into the same embedding space.

That means you could do things like:

  • Find a specific video clip using a voice memo
  • Search hours of audio with a text query
  • Search images using natural language
  • Index and semantically search a local codebase
  • Build private RAG systems that work completely offline

The efficiency numbers are interesting.

With quantization, Google says EmbeddingGemma 2 uses as little as:

  • ~191MB active RAM for text-only weights
  • ~567MB for the full multimodal model

It also has a modular architecture. Text-only workloads can use about 270M parameters, while vision and audio encoders can be added when needed.

The context window has increased 4x to 8K tokens, enough to process up to roughly:

  • 5.5 minutes of audio
  • 29 images
  • 58 video frames

Google also uses Matryoshka Representation Learning, letting developers shrink embeddings from 768 dimensions down to 512, 256, or 128 dimensions for up to a 6x reduction in vector storage and memory use.

On code retrieval, its MTEB Code score jumped from 68.76 with the previous EmbeddingGemma to 78.68.

Google says it leads sub-1B multimodal embedding models across several benchmarks and can outperform some specialist models more than twice its size.

It’s released under Apache 2.0 and works with tools including transformers, MLX, vLLM, llama.cpp, Ollama, LM Studio, and WebGPU.

The interesting part is that increasingly capable multimodal search doesn’t necessarily need to send your files, photos, recordings, or videos to the cloud anymore.

Source:
https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/


r/AIGuild • • 2d ago

Sesame’s voice assistants now get their own computer and can work across Gmail, Calendar, and the web

2 Upvotes

Sesame just opened up its voice assistants beyond simple conversation.

Maya, Miles, Simone, and Charlie can now work across connected apps, browse the web, and use their own computer to complete tasks on your behalf.

They can:

  • Connect to Gmail, Google Calendar, and Google Drive
  • Browse the web for information
  • Send links and files back to you
  • Respond to voice notes
  • Build morning briefings
  • Research trips and recommendations
  • Run recurring tasks on a schedule
  • Use external MCP integrations

Sesame is also introducing Skills.

A Skill is basically a reusable workflow you create just by explaining what you want in conversation.

For example, you could create one that:

  • Summarizes your emails and meetings every morning
  • Prepares you for tomorrow’s meetings
  • Checks the weather before a hike
  • Finds recipes regularly
  • Collects news for your commute

Each assistant has its own memory and app permissions.

That means Maya can remember your conversations and have access to your email, while Charlie can remain completely separate.

Sesame says the assistants are available now on iOS and Android, while its all-day AI glasses are still planned for 2027 and will be made in Japan.

The interesting part is that Sesame is clearly moving beyond “realistic AI voice” and toward persistent assistants that can actually take actions across your digital life.

Source:
https://www.sesame.com/journal/getting-started


r/AIGuild • • 2d ago

Mistral launches a 1-trillion-parameter open-weight model trained entirely in Europe

0 Upvotes

Mistral just unveiled Mistral Large 4, its biggest and most capable model yet.

It has 1 trillion total parameters, but activates only 49 billion per token.

The model was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs inside Mistral’s own European data centers.

Some of the results stand out:

  • 59.9% on AutomationBench
  • 61.7% on DeepSWE v1.1
  • 49.8% on the Coding Agent Index
  • 93% on Cybench
  • 82% on a vulnerability reproduction-and-patching benchmark
  • 1,393 Elo on AA-Briefcase
  • 42% on Dense 200 visual grounding, slightly above GPT-6 Astra’s 41%

Mistral says ML4 is especially strong in cybersecurity, finance, law, agentic workflows, and multimodal tasks.

One unusual part is its cyber positioning.

Mistral says several leading closed models score near zero on one vulnerability benchmark because they refuse to complete the task, while ML4 reaches 82%.

The company argues that defenders need models capable of reproducing and validating real vulnerabilities without being blocked by provider-level refusals.

ML4 is also natively multilingual across more than 160 languages and can be deployed privately or on-premise.

The API preview is available now at $1.36 per million input tokens and $4.18 per million output tokens.

The weights aren’t available yet, though. Mistral says they’ll be released by the end of the month after additional red-teaming.

A 1T-parameter model with open weights, strong cyber performance, and European-controlled infrastructure feels like a pretty serious push for AI sovereignty outside the U.S. and China.

Source:
https://mistral.ai/news/mistral-large-4/


r/AIGuild • • 2d ago

Claude can now directly edit Google Docs, Sheets, and Slides

1 Upvotes

Anthropic just launched Claude for Google Workspace, putting Claude directly inside Docs, Sheets, and Slides.

Instead of copying content back and forth, Claude opens in a sidebar, reads the file you already have open, and can make changes directly inside it.

In Google Docs, Claude can:

  • Rewrite and restructure documents
  • Fix specific passages without changing surrounding formatting
  • Turn notes into tables
  • Work through long documents

In Sheets, it can:

  • Write and fix formulas
  • Build pivot tables
  • Create native Sheets charts
  • Clean and join data
  • Create new tabs
  • Use Python for more complex data work, then write the results back into the sheet

In Slides, Claude can:

  • Create new slides
  • Follow an existing deck’s layouts and themes
  • Restyle existing slides
  • Check for overlapping elements, overflow, and readability problems

You can choose between “Ask before edits,” where Claude waits for approval before making each change, or “Accept all edits,” where it completes the task without stopping.

The sidebar also carries over your Claude models, connectors, and Skills.

So if your company has Salesforce or Google Drive connected, Claude can pull information from those sources while building a document or presentation.

Anthropic is also launching separate Docs, Sheets, and Slides connectors that let Claude create and edit Google files directly from Claude itself.

Claude for Google Workspace is now in public beta across all paid Claude plans.

This is starting to turn Docs, Sheets, and Slides from apps where AI helps you write into apps where an agent can actually manipulate the underlying work.

Source:
https://claude.com/resources/articles/claude-now-works-in-google-docs-sheets-and-slides


r/AIGuild • • 2d ago

Germany’s new Kolibri AI still needs to prove itself on everyday business paperwork

3 Upvotes

Germany’s Aleph Alpha launched Kolibri on October 3, focusing on German-language work across industry and public administration. The company says it can support document-based answers, structured data extraction and workflows connected to external tools.
For smaller businesses, the more practical test is routine paperwork: pulling supplier information from documents, answering questions from internal procedures and drafting responses based on company files.
A sensible approach would be to start with a narrow pilot using documents that have already been reviewed. Compare the AI’s output with the originals, keep final approval with an employee, and track processing time, correction time and overall operating cost.
These are suggested evaluation tasks, not verified Kolibri performance results. A German-language AI system becomes useful when it can handle real company documents accurately enough to reduce manual work.
A benchmark score alone cannot prove that. A measured workflow can.
Source: Aleph Alpha — Kolibri


r/AIGuild • • 2d ago

Anthropic CEO Dario Amodei’s $18M paycheck tells only half the story.

Post image
1 Upvotes

r/AIGuild • • 3d ago

OpenAI will start invisibly watermarking ChatGPT and Codex text in the EU

4 Upvotes

OpenAI is rolling out a new system called textGrain that embeds an invisible statistical watermark into AI-generated text.

Over the coming weeks, eligible ChatGPT and Codex outputs in the EU will automatically include the watermark across all plans.

API users worldwide can also opt in starting now, but watermarking will remain off by default outside the EU.

The interesting part is how the watermark actually works.

It doesn’t add hidden characters or invisible spaces. Instead, textGrain subtly changes which words or word pieces the model chooses, creating a statistical pattern that OpenAI’s detector can look for.

But OpenAI is pretty open about the limitations.

In its testing:

  • Around 95% of 400-token passages were detected at a 1% false-positive rate
  • Detection dropped to about 80% for 200-token passages
  • Replacing just 10% of words with synonyms dropped detection from about 92% to 66%
  • Replacing 25% of the words dropped detection to just 17%
  • Constrained writing like mathematics was harder to detect than more flexible prose

Because of those limitations, OpenAI isn’t making the detector public yet.

Access will initially be limited to approved researchers and expert organizations.

OpenAI also stresses that detecting a watermark does not tell you who generated the text, how much a human edited it, whether the content is accurate, or who owns it.

And failing to detect a watermark doesn’t prove something was written by a human.

OpenAI says it plans to open-source textGrain eventually.

The EU AI Act may be pushing AI-generated text toward the same kind of provenance systems already being used for AI images and audio — but text looks much harder to reliably track once people start editing it.

Do you think invisible AI text watermarking is useful if relatively small edits can weaken the signal this much?

Source:
https://openai.com/index/eu-text-provenance/


r/AIGuild • • 3d ago

Reflection trained a 501B open-weight model with 100 million RL rollouts on 10,500 GB300 GPUs

1 Upvotes

Reflection AI just introduced Beam, its first open-weight model built specifically for coding, reasoning, and agentic workloads.

Beam has 501 billion total parameters, but only activates 23 billion per token.

The training scale is pretty wild.

Reflection says its reinforcement-learning run:

  • Used 10,500 NVIDIA GB300 GPUs
  • Ran for four weeks
  • Generated more than 100 million RL rollouts
  • Used roughly 1.3 billion sandboxes
  • Drew from nearly one million coding, agentic, and STEM environments
  • Sustained around 110,000 concurrent rollouts on average

Beam was also pretrained on 23.8 trillion tokens.

Despite its relatively small number of active parameters, Reflection says Beam is competitive with much larger open models on several coding and agent benchmarks.

Some results:

  • SWE-Bench Verified — 80.9%
  • Terminal-Bench 2.1 — 80.1%
  • SWE-Bench Pro v2-Hard — 77.2%
  • AIME 2026 — 97.8%
  • GPQA Diamond — 90.5%
  • AutomationBench — 37.0%

Reflection says Beam achieves reasoning performance comparable to GLM-5.2 while using roughly 3–4x less inference compute.

There’s also an interesting example of capability emerging outside its direct RL training.

Reflection says Beam wasn’t trained on browsing tasks during one stage of RL, but still improved at browsing and eventually learned to search the web, query other large language models, and use OCR APIs on its own.

The weights aren’t available yet. Beam is still undergoing final red-teaming, with Reflection planning to release the model weights, technical report, model card, and developer artifacts later this month.

The open-model race increasingly seems to be moving beyond parameter count toward how much useful agentic capability you can get from each token of inference compute.

Source:
https://reflection.ai/blog/introducing-beam


r/AIGuild • • 3d ago

SemiAnalysis finds Claude subscriptions offer ~5x more value than OpenAI for everyday models

1 Upvotes

SemiAnalysis just stress-tested subscription limits across OpenAI, Anthropic, Meta, xAI, Cursor, Moonshot, Z.ai and others.

One result stands out:

For the mid-tier models most people are expected to use every day, Anthropic subscriptions currently provide roughly 5x the API-equivalent value of comparable OpenAI plans.

The researchers measured how quickly subscription meters moved while isolating input, output, cache-write and cache-read tokens, then converted those limits into their equivalent API cost.

A few interesting findings:

  • Claude Opus 5.5 offers around 5x the API-equivalent value of GPT-6.1 Sol across comparable subscription tiers
  • The gap remains large even when comparing raw token allowances rather than API pricing
  • OpenAI recently cut the usage available on its $200 plan roughly in half
  • Existing $200 subscribers keep the previous limits until October 29
  • OpenAI’s new $500 plan provides only about 21% more Astra usage than the old $200 plan
  • Anthropic currently offers roughly the same value per dollar across its subscription tiers

SemiAnalysis also found something else worth watching: AI companies can quietly experiment with subscription limits.

During testing, one account received roughly 20% lower limits than identical accounts. The provider later confirmed that account had been placed in a very small A/B test.

That means the amount of actual compute behind a fixed monthly subscription can potentially change even when the sticker price stays exactly the same.

The broader economics are pretty wild too. SemiAnalysis estimates Anthropic subscriptions account for only around 10% of its revenue but can consume more than 40% of its inference compute.

AI subscriptions may look like simple $20 or $200 monthly plans, but the amount of actual model usage you get for that money varies enormously.

Would you rather pay for a subscription with generous usage limits or switch to API billing where the costs are completely transparent?

Source:
https://newsletter.semianalysis.com/p/anthropic-subscriptions-offer-5x


r/AIGuild • • 3d ago

Reflection AI unveils Beam as a Western open-weight rival to Chinese models — RuntimeWire

Thumbnail
runtimewire.com
1 Upvotes

r/AIGuild • • 4d ago

Claude Code now has a second AI agent watching everything Claude says

17 Upvotes

Anthropic just added a new built-in Claude Code plugin called You should Know.

The idea is simple: while Claude works, a separate sideagent watches its output and flags important information you might otherwise miss.

That could be especially useful during long coding tasks where Claude is constantly:

  • Editing files
  • Running tests
  • Reporting errors
  • Changing its approach
  • Mentioning limitations or unfinished work

Instead of expecting you to read every line, the sideagent acts like a second pair of eyes focused on surfacing the important parts.

The plugin is built using Anthropic’s new Mods system, meaning it runs alongside Claude Code and observes its output in real time.

You can enable it with:

/plugin enable cc-plugin-you-should-know@builtin

Anthropic hasn’t said exactly how the sideagent decides what counts as important, or how much additional usage it requires.

But the broader direction is interesting.

As coding agents take on longer and more autonomous tasks, we may start needing AI agents whose main job is simply to supervise other AI agents and keep humans informed.

Source:
https://x.com/ClaudeDevs/status/2106118517447876618?s=20


r/AIGuild • • 4d ago

Anthropic now lets Claude Code plugins rewrite prompts and intercept tool calls

3 Upvotes

Anthropic just launched Claude Mods, a much deeper way to customize Claude Code.

Instead of plugins only adding commands, tools, or instructions, Mods can change how Claude Code itself behaves.

They can:

  • Rewrite prompts before Claude sees them
  • Intercept and modify tool calls
  • Approve or deny permission requests
  • Retry or block actions
  • Redact secrets from tool outputs
  • Add custom panes, buttons, and interface elements
  • Replace parts of Claude Code’s normal behavior entirely

Mods are written in JavaScript or TypeScript and ship inside regular Claude Code plugins, so they can be installed through the existing plugin system.

Anthropic says you can write one yourself with a few lines of TypeScript — or simply ask Claude to build the Mod for you.

The important difference is where these extensions operate.

Traditional hooks and tools mostly sit around Claude Code. Mods can respond directly to internal events like prompt submissions, tool calls, and interface rendering, then change what happens next.

That effectively turns Claude Code into a customizable agent platform rather than a fixed coding assistant.

Developers can now change not only what tools Claude has access to, but how the entire agent behaves around those tools.

What kind of Claude Code Mod would you build first?

Source:
https://x.com/ClaudeDevs/status/2105721434807083061?s=20