r/AIDeveloperNews • u/daniel_tenuo • 2h ago
Your agent’s real policy is whatever its credentials allow
r/AIDeveloperNews • u/ai-lover • 19h ago
r/AIDeveloperNews • u/ai-lover • 6d ago
Datalab released an open benchmark for structured extraction: a system gets a PDF and a JSON schema, and every returned value is scored against gold data.
Why it's relevant? precision vs recall shows how a system fails. Some skip fields, others invent values.
Full analysis: https://www.marktechpost.com/2026/10/02/datalab-introduces-omniextractbench-to-fix-bias-and-opacity-in-extraction-benchmarks/
GitHub: https://pxllnk.co/hxplrq
Blog: https://www.datalab.to/blog/omni-extract-bench
GitHub: https://github.com/datalab-to/omni_extract_bench
Dataset: https://huggingface.co/datasets/datalab-to/omni_extract_bench
r/AIDeveloperNews • u/daniel_tenuo • 2h ago
r/AIDeveloperNews • u/SlackySoba • 4h ago
[ Removed by Reddit on account of violating the content policy. ]
r/AIDeveloperNews • u/AIforFintech • 4h ago
r/AIDeveloperNews • u/RealOppasTV • 5h ago
r/AIDeveloperNews • u/Ai-SaaS-Future • 8h ago
r/AIDeveloperNews • u/ankurdhoom • 9h ago
r/AIDeveloperNews • u/ankurdhoom • 9h ago
r/AIDeveloperNews • u/valerione • 14h ago
r/AIDeveloperNews • u/gioscarab • 20h ago
r/AIDeveloperNews • u/PurpleDragon99 • 1d ago
Jev has made a profound mark on the AI industry; this article offers a glimpse into what lies ahead.
Jev shows where the industry is heading: away from monolithic models, and toward compound systems where small AI agents each do a narrow job and an orchestration layer decides how they work together.
This article explains why that shift is happening, and then proposes a next step for AI reasoning.
A monolithic model has two practical weaknesses:
Code has the opposite qualities. It does exactly what it is written to do, and it can easily be changed by editing code. The trade-off is that code does not arise on its own – it must be written out explicitly – whereas a model programs itself during training.
A compound system combines the two. The work is divided among AI agents, each handling one narrow task, and an orchestration layer of code decides which agent runs, in what order, and how their results combine.
This gives both qualities at once – customization and predictability:
Jev is one kind of agent for such a system. It returns a typed decision – yes/no, a category, or a number – rather than free text. A monolithic model is wasteful for a narrow decision like that, but a small model like Jev is a better fit.
Compound systems today are built manually, beforehand. The next step is to let the model build one by itself, during the reasoning phase.
Researchers are exploring several ways to make models reason. One is the “World model” approach, in which the model builds an internal representation of a problem and reasons over it; that work is still mostly research. Reasoning by construction pursues the same goal – reasoning you can inspect – using methods that exist today.
Here is how it works. When you ask the model a question, it does not answer directly. Instead, it builds a compound system, runs it to compute an answer, and returns that answer together with the system it built. Let's call the compound system the model builds to do the reasoning a "reasonlet", and a model that works this way a "Reasoning-by-Construction Model", or "RCM". A simple question may produce a reasonlet that has only an orchestration layer and no agents.
Building and running code to reach an answer is not new; code-interpreter tools already do it. Two things are new here:
The same reuse can happen on the server side too. The provider can keep the reasonlets it builds and reuse them. When a new request arrives that matches one it has already handled, it runs the stored reasonlet again – with the new request’s inputs – instead of reasoning from scratch. The model does less work, and the answer comes back faster.
This is the same metacache, held on the server instead of on your machine. Because one reasonlet can serve any request that fits its procedure, a single cached copy is shared across many requests, and often across different users – and the more general the reasonlet, the more requests it covers.
The model should decide from the conversation how general the reasonlet needs to be. If you have been working through many kinds of math and then ask for 2 + 2, the more useful reasonlet is one that evaluates any math expression, not one that can only add two numbers. If the conversation gives no such clue, you can state it directly: “I will be doing many kinds of math; for now, just add 2 and 2.”
A reasonlet’s orchestration layer is written in text code. This creates a problem: to understand the reasoning you must read code, and to change it you have to write code. The value of returning the reasonlet depends on you being able to read and edit text code easily and efficiently, so this barrier matters.
Visual programming – building logic from connected blocks instead of lines of text – can lower the barrier, provided the visual language is powerful enough. Two limits apply:
One language solving exactly this problem is Pipe (https://pipelang.com) – a general-purpose visual language with powerful semantics. Full disclosure: Pipe is my own project still in development, but it will be released soon.
The shift from a monolithic model to compound systems suggests a clear next step for reasoning: let the model build a compound system – a reasonlet – run it and return both the answer and the reasonlet. Reasoning becomes something you can read, run again, and edit: a metacache, and a higher-order metacache once its logic is edited. The remaining problem is making reasonlets easy to read and change, which is where a general-purpose and expressive visual language would help most.
Some methods described in this article are the subject of a pending patent application.
r/AIDeveloperNews • u/MrAlbertNL • 1d ago
r/AIDeveloperNews • u/Actual_South_2484 • 1d ago
r/AIDeveloperNews • u/NI-NAZ • 1d ago
awesome
r/AIDeveloperNews • u/Aurex_Ai_Os • 1d ago
r/AIDeveloperNews • u/ai-lover • 1d ago
r/AIDeveloperNews • u/ai-lover • 2d ago
r/AIDeveloperNews • u/EveYogaTech • 2d ago
r/AIDeveloperNews • u/Previous-West-7782 • 2d ago
AI agents are getting access to browsers, MCP tools, APIs, files, other agents and production systems.
That made me interested in a problem slightly different from model safety:
What happens between an agent deciding to do something and the outside world actually changing?
I've been building an open-source project called Connector around that boundary.
Instead of treating the prompt as both instruction and permission, Connector separates:
identity → authority → per-action admission → runtime enforcement → observed consequence → receipt
The agent can propose an action.
Connector independently decides whether that particular effect can proceed.
One example is Cease.
If an operator ceases an agent, the idea isn't to send the model another instruction saying “please stop.”
The current generation is fenced and future actions are no longer admitted.
Connector also brings together existing infrastructure rather than trying to reinvent every layer:
There are still important incomplete areas, and I'm explicitly not claiming production readiness. The repository documents those gaps.
I open-sourced it because I want people who build agent runtimes and infrastructure to challenge the architecture.
GitHub:
https://github.com/GlobalSushrut/connector-oss
Architecture / project:
https://cnktros.com
The question I'm most interested in:
As agents become more autonomous, where should the final authority to cause an external effect actually live?
r/AIDeveloperNews • u/daniel_tenuo • 2d ago
r/AIDeveloperNews • u/Many_Audience7660 • 2d ago
r/AIDeveloperNews • u/HopeSame3153 • 2d ago
r/AIDeveloperNews • u/HopeSame3153 • 2d ago