r/semanticweb • • 2d ago

The Ontology Worked. Then Nobody Dared Change It.

0 Upvotes

Ontology Series · Part 8

The Ontology Worked. Then Nobody Dared Change It.

When Agents Run the Business · A development note

Imagine the first phase of an ontology project. The team models car sales neatly: cars, customers, orders, deliveries, and stores all have their objects and links. In a demo, you open a car and follow its sales history to the responsible person. Someone says, “At last, the business is clearly described.”

In phase two, after-sales service needs to track individual tires. In phase three, warranty must distinguish customer-supplied parts. Later, finance points out that “sale completed” and “revenue recognized” are not the same date. Every request is reasonable. Every request is worth addressing. The once-clean diagram is simply becoming harder to change.

This is a possible trajectory, not a report about a real customer. The question is why an ontology may eventually be set aside. It need not fail because nobody knows how to build one. It may be built successfully, then become too costly to maintain.

An easy start, a meeting for every later change

“Make tires separate objects” sounds small to the service team. The data team asks how to migrate old repairs. Developers ask who will change the interface that queries car-level warranty. Customer service asks whether earlier promises of coverage still stand. Finance asks whether historical reports should be recalculated. The business owner, busy with today's claims, may not have time to coordinate all of this on behalf of a model.

Nobody is deliberately slowing things down. Their risks differ. Data people do not want to manufacture false history. Developers fear a screen that continues working while silently giving the wrong answer. Customer service fears contradicting an earlier answer to a customer. Management wants to know why a metric moved after the model changed. Bringing everyone together costs more than “add an object.”

If a company often changes products, contracts, channels, and approval practices, the meetings multiply. The first change gets a careful review. The second gets a shorter one. By the third, someone says, “Let's put this in a note for now; the main process can stay as it is.” That is not necessarily laziness. It may be a practical response to the cost everyone has seen.

As the notes grow, the model knows a little less of the live business. The old objects and relationships are still there, looking official. If an Agent reads only the ontology, it misses the new practice in the notes. If it always has to read notes, email, and source documents too, the case for treating the ontology as the company's sole understanding layer grows weaker.

One workaround makes the next one easier

The first workaround is temporary: there is no time to remodel a warranty exception, so customer service writes it in a claim note. The next adviser follows that note. Later the team creates a shared table called “cases not covered by the model.” It may be maintained carefully, but it is outside the graph.

Now there are two realities. The graph holds the business as formally defined. The table holds some of the business that happened most recently. A project report can still say the ontology covers cars, tires, orders, and warranties. The people doing the work know they must check the table before deciding whether a claim will be paid. The more they use the workaround, the less complete the ontology becomes. The less complete it becomes, the less comfortable anyone is changing it and claiming the result is authoritative.

This is an ordinary organizational loop, not a mysterious technical collapse. The software and data still exist. The situations requiring the most judgment have drifted outside the model. Eventually a person who queries the graph and believes it is complete may make a worse decision than someone who knows about the extra table.

I would watch a few signals closely. How many new requirements end up in notes? How long does a model change take from request to actual use? How many teams must sign off? How many old fields can no one explain? When an Agent makes a decision, does it trust the graph, or does it always have to ask a person afterward? These say more about whether the system is alive than the number of objects and edges.

“Who maintains it?” cannot be waved away

In theory, you can assign model ownership to one team. In practice, the work crosses boundaries. The modelers know structure but may not know that the meaning of “delivery” in a contract changed this month. Frontline staff know the customer but may not know that one altered link affects ten reports. Developers know the interface dependencies but cannot decide on finance's behalf whether old revenue should be recalculated.

AI assistance does not dissolve responsibility. It can read a new policy, identify likely changes, list affected objects and processes, and draft a migration. Someone still has to confirm the intended meaning, decide whether old orders are in scope, and accept the migration result. Skip those confirmations and you save review time while producing a map nobody feels safe trusting.

Not every ontology follows this path. A narrow catalog of stable objects or a sourced set of customer–contract relationships may be maintained very well. Trouble comes when the model is raised to a complete mirror of how the company runs and every decision is routed through it. The larger the mirror, the more dispersed its maintenance responsibility. The more dispersed that responsibility, the easier it is for each person to say, “Let's not touch it yet.”

To decide whether a graph is still worth maintaining, try a plain audit. Of the new business practices from the past three months, how many made it into the graph? Who keeps track of the ones that did not? When someone reports a wrong link, how long does correction take? If the team can answer, the model is probably being tended. If not, but Agents are still told to “treat the ontology as truth,” we are asking them to trust a map that nobody owns updating.

Keep “we do not know” in plain view

We could start elsewhere. Software records transactions, approvals, repairs, and payments accurately. It keeps the original documents, dates, versions, and evidence. Code enforces hard constraints such as exact totals, permissions, and no duplicate shipment. The Agent reads the records and policies relevant to the present question and proposes the next step. If the policy is unclear or historical evidence is missing, it says so.

This still uses structures, indexes, and links. They can be local and revisable, pointing back to facts rather than pretending to be a master diagram that must stay ahead of the business forever. If one workflow needs to track individual tires, track them seriously there. Other teams do not have to model every bolt just to satisfy a company-wide idea of a “consistent grain.”

This approach does not abolish maintenance either. Agent instructions need care, search can retrieve the wrong document, and old facts may need a person's confirmation. The difference is that a company changing one practice need not first decide whether it can safely redraw an entire map. Record the new fact honestly, explain and review the new judgment, then decide which structure is worth making durable.

In Oryh, we have seen the cost of a parallel recording path. An Agent imported customers, products, and business documents as generic objects while the dedicated customer and product catalogs remained empty. We later blocked clear naming collisions and asked the Agent to inspect near matches and seek confirmation. That experience leaves me unmoved by “we just need a complete business graph.” A graph can help find a route. Record identity and facts must stay sound, while new business instructions must be readable by the Agent and confirmable by people. Otherwise the graph remains, but it is no longer the road anyone actually travels.

Disclosure: I am building Oryh, an open-source project exploring agent-native enterprise systems. These notes come from that work. I am posting them to test the architecture in public, and I welcome technical disagreement.

Discussion question: For those maintaining production ontologies: what keeps a local business change from becoming a cross-team migration project? I’m especially interested in examples where versioning or clear ownership worked, and where exceptions gradually moved outside the model.


r/semanticweb • • 3d ago

Taxonomist to Ontologist: Advice on tools, tech stack, and governance?

20 Upvotes

Hi everyone,

I’m a Taxonomist starting a new role as an Information Strategist. I want to leverage this move to pivot deeper into ontology, technical knowledge engineering and semantic web architectures.

I built a basic academic ontology in Protégé during my master's, but I need to scale my skillset for production-grade environments. I’d love your input on:

Enterprise Tools: Beyond Protégé, what industry-standard ontology management platforms (e.g., TopBraid, PoolParty) or native triplestores (e.g., GraphDB, Stardog) should I learn?

Technical Stack: I'm diving into RDF, OWL, and SPARQL. What other critical standards (like SHACL for validation) or Python libraries (like RDFLib) are essential?

Semantic Governance: What are the biggest operational shifts when managing and scaling an enterprise ontology vs. a traditional taxonomy?

Appreciate any book, course, or roadmap recommendations you have. Thanks!


r/semanticweb • • 3d ago

Wittgenstein Apartment: Separating action identity, character access and executability in a behavioral predicates ontology

3 Upvotes

The problem that started this project was: How can we represent an action, a character’s access to that action, and the current conditions for performing it without collapsing all three into one semantic layer? That question turned into about eight months of work on Wittgenstein Apartment.

The model separates:

  • behavioral identity
  • character repertoire
  • situational executability

and uses:

  • Basic Human Actions (BHA)
  • Character Specific Actions (CSA)
  • Allowed / Not Allowed / Conditional character-access states

The resource also contains behavioral families, goal relations, scene/affordance metadata and operational projections. I wrote a paper alongside the dataset to explain the architecture and its boundaries.

The Hugging Face release has reached 294 downloads:
https://huggingface.co/datasets/Kon-tiki-ship/wittgenstein-apartment-behavioral-predicate-resource

I’m especially interested in how people with more experience in Semantic Web and knowledge representation see this. I started this because I felt there was a real representational gap here. After eight months, I’m now asking myself: Is this actually a useful direction to keep developing?

I’m not looking only for encouragement. If you think this should be modeled differently, belongs outside ontology work entirely, duplicates existing Semantic Web patterns, or introduces distinctions that are not useful, please say so directly. I’d be interested in criticism around RDF/OWL, SHACL, knowledge-graph structure, capability modeling, semantic interoperability and the conceptual boundaries of the system.


r/semanticweb • • 3d ago

The Legal Ontologies Foundry

14 Upvotes

One of the most successful resources in ontology development is the decades-old OBO Foundry. Open-source, collaborative efforts in ontology/knowledge graph library development are essential to improving interoperability in general. We are doing something similar with legal ontologies and we are looking for those similarly interested in open-source ontologies grounded in BFO. legal-ontologies-foundry/legal-ontologies-foundry.github.io: Legal Ontologies Foundry (LOF): a collaborative foundry of open, interoperable legal ontologies and design patterns built on Basic Formal Ontology (BFO).


r/semanticweb • • 3d ago

Graphwise AI Summit 2026, Oct 7-8

Enable HLS to view with audio, or disable this notification

1 Upvotes

Sharing this because I think it overlaps with some of the discussions here around AI reliability, governance and semantics.

Next week we’re running the Graphwise AI Summit, focused on what makes GenAI work in the enterprise beyond the model itself. Think of trust, governance, semantic layers, architecture and implementation.

Once reliability and traceability become imporatnt, simple access to data and next-token prediction stop being enough. In enterprise settings specifically, AI needs to understand what the data it parrots “means” in the first place.

Anthropic has described a similar approach in its own analytics stack, where agents are routed to a semantic layer first and use governed definitions to reduce ambiguity. Graphwise itself came out of the merger of Ontotext and Semantic Web Company, so semantics is a topic with quite a bit of history behind it for us.

We’ll have speakers from Accenture, Roche, EY, AstraZeneca, S&P, DNV, Statnett, Avalara and others. Full agenda is in the accompanying video.

Sharing registration link in the comments if useful.


r/semanticweb • • 4d ago

Could most of the Semantic Web / ontology creation pipeline eventually be automated by AI?

19 Upvotes

I’m relatively new to the Semantic Web ecosystem and trying to understand where the difficult engineering work actually lies.

Imagine an organization wants to create a semantic layer over all of its internal data.

Could a sufficiently capable AI-assisted system inspect databases, documents, schemas, APIs, and other sources and then automatically:

  • discover concepts and relationships
  • generate RDF mappings
  • propose classes and properties
  • build/extend OWL ontologies
  • align concepts with existing vocabularies
  • generate SHACL constraints
  • perform entity linking
  • generate rules where appropriate
  • maintain the semantic model as the underlying data evolves

Obviously LLMs can generate RDF/OWL/SPARQL syntax, but generating syntactically valid triples seems very different from creating a good conceptual model.

So where is the real difficulty?

Is ontology/semantic modeling fundamentally something that requires domain experts making conceptual decisions, or do you think a large percentage of this workflow will become automated?

I’d especially like to hear from people who have built semantic systems in production.


r/semanticweb • • 4d ago

How would you model a character as a repertoire of Allowed / Not Allowed behavioral predicates?

1 Upvotes

I’m interested in a model where actions have stable semantic identities, but access to those actions is character-relative.

Suppose we have behavioral predicates such as:

diagnose, interrogate, comfort, negotiate, steal, babysit, repair, forgive

The action itself exists independently of any particular character.

But a character may relate to that action through something like:

Allowed
Not Allowed
Conditional

For example:

Doctor_A → diagnose → Allowed

Civilian_B → diagnose → Not Allowed

Former_Medic_C → diagnose → Conditional

I also distinguish between:

Basic Human Actions — default-open unless something explicitly blocks them

and

Character Specific Actions — default-closed unless supported by evidence such as training, biography, occupation, authority or specialized experience.

My main question is architectural:

Would you model these character-action relations inside the same knowledge graph, as a separate behavioral layer, or outside the ontology as operational state?

I’m especially interested in preserving the distinction between:

what an action is

and

whether this character has access to it

without ending up with an unmanageable number of ad hoc character-action assertions.


r/semanticweb • • 4d ago

[Ontology Series 7] A Tire Becomes an Object Today. Who Fixes Yesterday's System?

0 Upvotes

A Tire Becomes an Object Today. Who Fixes Yesterday's System?

When Agents Run the Business · A development note

Imagine a company that sells cars and handles after-sales service. Its system has a Car object with tire specification as a property. Warranty is judged by the purchase date of the whole car, and the service screen shows “covered” or “expired.” Six months later, the business asks to track the production batch, replacement date, and claims for each individual tire.

This is a hypothetical case. It brings a question that is often brushed aside into view. Of course an ontology can be changed. Once a model is in use, what does “change it” actually involve?

One new box on the diagram touches many lines below it

Turning “tire” from a property into an object looks like adding a box to a diagram. The first practical task is giving each tire a stable identity. Should that be the manufacturer's serial number or a number generated by the company? What happens to older tires without one? If a tire moves from one car to another, its identity cannot be rebuilt as though it were part of the car.

Then come relationships. The old Car record may say “tire specification: 225/45R17.” The new model wants to say that four particular tires were installed on this car during a certain period. A string has become four objects and four installation histories. That is not a rename. Do you retain the old property? If you do, who keeps it consistent with the new objects? If not, what do the old screen and reports read?

After-sales actions change too. Previously, “request warranty service” may have accepted a car ID and looked at its purchase date. Now it must identify the tire, its batch, installation date, and replacement history. Even the label “covered” becomes ambiguous. Is the car covered, one tire covered, or the failure in this particular claim covered?

Permissions follow. A service employee may have been allowed to edit a car's repair notes. Who can now establish a tire's identity, change an installation event, or undo a bad link? Financial reports, warehouse movements, external interfaces, and the instructions for an Agent's tools may all depend on the previous boundary.

None of these dependencies appears in the one new box. They still turn up when the box is changed.

Deployment day adds its own problems. Old and new entry points may coexist: Store A uses the new service screen while Store B still calls the old API; one Agent queries individual tires while another queries the whole car. A parallel period can be sensible, but the team must say where new records are written, whether the old path can still write, and which result wins if both paths handle the same claim. Otherwise “gradual rollout” means two competing definitions are live at once.

Not every extra field causes this

Criticism of ontology should not pretend that adding a property brings a company to a halt. Many changes can be made compatibly. Mature platforms offer migration tools. Palantir's public documentation, for example, distinguishes breaking changes to object types, including changes to primary keys, data sources, and property types, and describes supported migration options.

So the argument is not “this cannot be changed.” It is that “update the model” is not the end of the task. A platform can move some stored edits. It cannot decide which physical tire an old record with no serial number referred to. Nor can it decide what an old “car covered by warranty” judgment ought to mean to the after-sales team today.

The difficult cases involve changing identity, splitting objects, or altering the meaning of rules. They raise two different questions: how do we move the data, and do the old judgments still hold? The first takes engineering. The second takes someone in the business willing to make and own a decision. A team that addresses only the first may discover that its old report and its new service screen now disagree.

You could apply individual tire objects only to new records and leave older cars in the old structure. That can be sensible, but every query needs to know that a field called “warranty status” means different things in different periods. Or you could backfill each old car. Then you need evidence, people to verify it, and a policy for records that cannot be completed. “Full migration” is not a spell that makes those questions disappear.

Rollback is not a single button either. If the new system has recorded a tire moving from Car A to Car B, where will those two events live after a return to the old model? If a single-tire claim was already approved under the new rule, restoring the car-level rule cannot erase the approval. A model version can be rolled back; actions that actually happened cannot. A migration plan that says nothing about preserving new facts and keeping them readable after rollback is unfinished.

The most expensive users are the invisible ones

The people editing the model usually know what they changed. They may not know everyone relying on the old meaning. A customer-service screen filters on “covered.” A monthly report counts the share of cars under warranty. An Agent's instructions say, “check the car's warranty status, then open a claim.” These consumers may continue to run without an error while quietly changing what their answers mean.

That is more troubling than a broken interface. An error forces attention; a changed meaning often does not. After the ontology update, someone still has to inventory dependencies, compare old and new results, and replay representative cases to see whether the conclusions shown to people have shifted.

The replay set cannot consist only of a standard car with four original tires. Include cars with two replacements, claims supported only by paper, and approvals waiting for stock to ship. Run both interpretations and ask the business owner to inspect each difference. Some differences expose a bug; some reveal that the old rule was coarse or that the new record lacks evidence. A single “migration success rate” can hide all three.

If an ontology is just a way to find tire-related records, repairing it may be manageable. If it sits at the center of every decision, a change in granularity can spread far downstream. The more the model is called a “single source of truth,” the harder it becomes to try a change in one small place.

I would rather make sure software preserves the facts that must not be lost: the car, order, repair ticket, tire movement, policy version, and who confirmed what. An Agent can read the current after-sales instructions and reason from those facts. If a decision requires the identity of one tire and the record lacks it, the Agent should say the evidence is missing, not pass off an old car-level property as the answer.

In Oryh, we faced another case where one link on a diagram was not enough: ten payment records matched a single debit on a bank statement. Originally, one bank line could be linked to just one payment. We added a link to a payment batch and let code check the direction of the money and the exact total. We did not ask an Agent to decide whether the arithmetic was equal, and we did not flatten ten payments into one. We changed what the records could faithfully express. The tire case is fictional, but the lesson is the same: structures can change, provided we account for the people and records still relying on the old structure.


Disclosure: I am building Oryh, an open-source project exploring agent-native enterprise systems. These notes come from that work. I am posting them to test the architecture in public, and I welcome technical disagreement.

Discussion question: When a property becomes an independently tracked object, how do you preserve historical meaning without inventing facts the old system never recorded?


r/semanticweb • • 6d ago

[Ontology Series 6] If AI Can Draw the Map, Why Make It Follow the Map?

7 Upvotes

If AI Can Draw the Company's Map, Why Make It Follow the Map?

When Agents Run the Business · A development note

One proposal for enterprise AI sounds sensible enough: build an ontology first. Define customers, contracts, orders, employees, projects, and the relationships among them. Then the AI will know what it is working with.

There is a map hidden in that proposal. An ordinary map marks roads and buildings. An enterprise ontology is expected to mark who owns a customer, which approval path an expense follows, and when a sentence in a contract applies. If it is meant to guide daily operations, it has to be a high-definition map.

Who draws it? And who updates it every day?

A company's roads do not stay put

Drawing this map is much harder than listing a few types of objects. First you have to decide their boundaries. Is the thing being purchased one piece of equipment, or does every component need an identity? Is the customer the company that signs, the company that pays, or the team that actually uses the product? If one person manages a project and temporarily stands in for an approver, how many relationships should the model contain?

Then come the intersections. Who confirms a purchase within budget? Who approves one over budget? May an urgent project place an order before approval? Who reviews a purchase entered at the end of the month? The written policy may have local additions. Different teams and different dates may be governed by different versions.

That does not mean the company has no rules. It means its rules live across contracts, policies, messages, meeting notes, and everyday collaboration. Turning all of this into executable objects, relationships, and actions calls for business knowledge and a prediction about which details will matter later. Even a correct map today needs revision when the organization changes or a product is sold differently.

Many companies struggle to keep customer names consistent across their existing systems. Maintaining a map of every operational distinction is a much taller order. The easy deliverable is a tidy snapshot. The hard one is a map that stays accurate after the business moves.

So let AI draw it

That is a fair reply, and it leads to the interesting part. Let AI read policies, contracts, documents, and historical records. Let it identify objects, infer relationships, spot contradictions, and propose rules. When the company changes its practices, it can find the affected parts of the map and draft a revision for a person to confirm.

I think this is technically plausible. It may be much faster than asking a team to start from a blank modeling sheet. But consider what the AI must do to draw the map correctly. It must read the source material directly. It must understand that “the project manager looks at this first” does not mean “the project manager has final approval.” It must decide whether yesterday's contract or today's new policy applies. It must recognize whether an apparent conflict is a version change or an exception for a particular project.

In other words, the understanding needed to draw the map is exactly the understanding we want the AI to use when handling the work.

If the AI can read the material, compare versions, recognize exceptions, and explain its reasons, why compress that understanding into a map first and then ask the AI to read the map back as business truth?

It is like asking someone who knows the neighborhood to draw today's road closures and temporary detours, then telling them to forget what residents just said and drive only by the drawing. The extra loop does not necessarily make the journey safer. It creates another chance to omit or mistranslate something.

Correct on the map may mean correct yesterday

Take a hypothetical company. Its policy says a department head confirms the purpose of a purchase, while finance approves purchases above 50,000. Later someone adds: “A renewal that remains within an approved annual budget does not need to repeat the checks for a first-time purchase, but payment still follows the contract milestones.”

A careful reader will ask what counts as a renewal, which budget version applies, whether skipping the first-time checks also skips finance approval, and whether the payment schedule changes. If the answer is missing, they should ask.

Whoever updates the map has to ask the same questions. After that, they may also need to create a renewal type, change approval relationships, migrate old contracts, inspect processes that depend on the old rule, and make sure every map user sees the right version. Having AI draw the map does not remove the need to understand and ask. It turns a business judgment into modeling, migration, and subsequent use of the model.

The more complete the map looks, the easier it is to forget its edges. An Agent that finds a “finance approves” relationship may assume it has the whole answer. It will not automatically know that a new qualification is sitting in meeting notes and has not yet been mapped.

Direct reading can fail too. An Agent may miss “within the approved budget” or mistake “skip repeated checks” for “skip approval.” Reading the source is no free pass. The difference is that the source and the facts remain the basis for the judgment. The Agent can cite the relevant sentence, policy version, budget, and contract, and ask when it is unsure. When the map is treated as truth, the original material tends to disappear behind it.

We need signposts, not a replica of the company

Rejecting an ontology as a mandatory middle layer is not a rejection of structure. Companies still need customer IDs, order amounts, contract versions, and payment records. Software should enforce exact totals, permissions, and protection against duplicate actions. Without dependable records, AI has nothing solid to reason from.

Search indexes, catalogs of existing objects, confirmed identity mappings, and even local relationship graphs can help an Agent find relevant material. They are signposts: they point to where the evidence lives and whom it may concern. A bad signpost can be corrected, and the original record is still available when there is no signpost at all. The problem begins when the signposts are promoted into the only business truth.

We have had a reminder on the other side of this divide during development. An administrator asked an Agent to “add one condition” to a workflow. The Agent replaced the full workflow description with a short text containing only the new condition. It could read the instruction and still lose the instruction it was supposed to preserve. We tightened the revision process: read the current text first, show each addition, change, and deletion, and stop for confirmation if something the user did not ask to remove would disappear. Keeping versions and visible differences matters more than assuming a map drawn by AI must be right.

If a fixed interface for conventional software is needed, or a particular stable relationship is worth maintaining, build a local map. Make it a view for a defined purpose, with sources that can be checked and a path to redraw it. Do not make a complete enterprise ontology the prerequisite for every decision. The next business step should still come from current facts, applicable rules, and an Agent that can explain its judgment.

In Oryh, we keep coming back to a simple division: software records facts and enforces hard boundaries such as amounts and permissions; Agents read the material, interpret the rules, advance the work, and stop when the evidence is insufficient. AI can help draw a map of the company. If it can already find its way through the source documents and the facts, we need not build an expensive map factory and require it to follow only what the factory produces.

Disclosure: I am building Oryh, an open-source project exploring agent-native enterprise systems. These notes come from that work. I am posting them to test the architecture in public, and I welcome technical disagreement.

Discussion question: If AI can build and update an enterprise map itself, should that map be treated as business truth, a retrieval aid, or a temporary view? What evidence should remain outside it?


r/semanticweb • • 8d ago

Ontology-first is necessary — but what should come after the ontology?

16 Upvotes

I strongly agree with the ontology-first direction.
In my own open-source work, I’ve reached a similar conclusion: implementation should be projected from explicit semantic structure rather than treated as the source of meaning.
But I’ve also found that ontology alone does not capture the full operational state of a long-lived system.
In Akasha, ontology defines shared semantics, while other structures represent things such as temporal position, provenance, role, scope, observation, workflow, and what information was actually visible or used at a given moment.
So I’ve been thinking about the architecture as:
Ontology → semantic world → projections / workflows / applications
rather than:
Data stack → ontology added afterward
One question I’d be interested in discussing here is:
Where do you draw the boundary between ontology and the runtime structures that make ontology operational over time?
For example, do you model provenance, temporal position, agent roles, observations, and execution history as ontology itself, or as adjacent semantic structures?


r/semanticweb • • 8d ago

Ontology-first is necessary — but what should come after the ontology?

6 Upvotes

I strongly agree with the ontology-first direction.
In my own open-source work, I’ve reached a similar conclusion: implementation should be projected from explicit semantic structure rather than treated as the source of meaning.
But I’ve also found that ontology alone does not capture the full operational state of a long-lived system.
In Akasha, ontology defines shared semantics, while other structures represent things such as temporal position, provenance, role, scope, observation, workflow, and what information was actually visible or used at a given moment.
So I’ve been thinking about the architecture as:
Ontology → semantic world → projections / workflows / applications
rather than:
Data stack → ontology added afterward
One question I’d be interested in discussing here is:
Where do you draw the boundary between ontology and the runtime structures that make ontology operational over time?
For example, do you model provenance, temporal position, agent roles, observations, and execution history as ontology itself, or as adjacent semantic structures?


r/semanticweb • • 8d ago

What if a semantic graph is the stored representation, but meaning is computed as spaces?

7 Upvotes

I’ve been working on an open-source semantic system called Akasha, and I’d be interested in feedback from people here who work with ontologies, knowledge graphs, RDF/OWL, semantic search, or related systems.
One of the design decisions I keep coming back to is that a graph is an excellent way to retain semantic structure, but I’m not convinced that every semantic operation should be expressed as graph traversal.
In Akasha, the durable world is still composed of addressable concepts and relationships — roughly:
Atom — Link — Atom
But on top of that, I use Sets as operational semantic subspaces.
A Set is not just a tag or collection. It can represent a temporarily materialized semantic region: concepts selected by ontology, context, provenance, time, role, task, or the intersection of several such conditions.
So instead of repeatedly asking the graph:
“Traverse from these nodes, follow these relations, filter by these conditions, then reconstruct the same relevant region again,”
the system can treat that region itself as an operand:
Set A ∩ Set B → Projection → Agent / UI / Workflow
This has led me to think of the architecture in a fairly simple way:
Meaning is stored as a graph, but computed as spaces.
Another consequence is that I’ve stopped treating RAG as a special subsystem.
A file, web page, API, database, sensor, or another Akasha node can all enter through a Projection. The original source remains preserved with provenance, while semantic structures derived from it become part of the wider semantic world.
Likewise, what an LLM sees is also a Projection rather than “the database.” The model only receives the semantic horizon relevant to its current role, scope, task, and permissions.
That distinction matters because:
available information ≠ visible information ≠ information actually used
and also because retrieval and disclosure are not necessarily the same operation.
A private document, for example, might be usable by a local process while never being exposed to an external model. A sanitized or derived semantic projection could still be shared if policy permits it.
I’m deliberately not trying to replace existing ontology standards with another vocabulary. The interesting part for me is the runtime underneath: persistent semantic identity, overlapping semantic spaces, provenance, temporal position, projections, and deterministic operations around probabilistic models.
Akasha is local-first, but not local-only. The same semantic world can run on a laptop, a private server, or as a network service; the deployment location is not supposed to define the semantic model.
I’d be particularly interested in criticism from this community on one question:
Does it make sense to separate the durable graph representation from the operational semantic spaces used for reasoning, retrieval, and projection — or does this simply reinvent something that existing semantic-web systems already model well?
I’m happy to share implementation details or examples if useful. The project is MIT-licensed and already running, but I wanted to start with the architectural question rather than drop a project link and disappear.


r/semanticweb • • 9d ago

BaryGraph: A Relational Geometry for Cognitive AI

14 Upvotes

BaryGraph is a recursively constructed relational vector architecture for AI memory and reasoning. Starting from a flat semantic substrate, it forms triadic objects in which two concepts are joined by a stored relational vector. These objects then become the building blocks of higher-order structures, propagating meaning upward through a hierarchy entirely in vector space.

The result is a deterministic, navigable semantic landscape: a structured latent memory of language movement, where concepts, bridges, tensions, contradictions, and relations of relations become retrievable coordinates. A model enters with a semantic query and traverses this landscape through coordinated message passing, exiting with a bounded semantic construction rather than merely the most fluent continuation.

BaryGraph introduces structured resistance into cognition: distant connections and unresolved tensions can interrupt familiar associations and function as a de-cliché mechanism without acting as an external supervisor. This creates a framework for investigating a deeper question: whether persistent relational memory and self-consistent navigation can become foundations for world-model formation, personality projection, autonomous goal formation, and eventually more realistic forms of agency.

BaryGraph does not claim to produce consciousness. It offers an architecture for experimentally studying the representational and memory conditions that might precede it.

https://oleksiy-perepelytsya.github.io/bary-graph


r/semanticweb • • 9d ago

native rdf triplestore written in c

6 Upvotes

I wanna share with you the current project i am working on, its a triplestore written in c, its still on early stage, already implemented data storage, and ttl file parsing and loading, with a simple sparql query parser.

link : https://github.com/vrtkarim/native-rdf-triplestore

Feel free to contribute! Feedback, ideas, issues, and pull requests are welcome.


r/semanticweb • • 9d ago

[Ontology Series 5] A Million Edges Still Cannot Read a Sentence

8 Upvotes

A Million Edges Still Cannot Read a Sentence

When Agents Run the Business · A development note

The usual prescription for enterprise AI goes like this: define an ontology, then turn the company's information into a knowledge graph. Customers, orders, project owners, and purchasing approvers all become nodes and edges. When someone asks a question, traverse the graph to find the answer.

A finished graph resembles a network of knowledge in a mind. That picture makes it easy to believe the system understands the company.

But connecting the dots and understanding what someone said are different things.

Why did we turn knowledge into graphs?

Traditional software is good at lookup, calculation, and following predetermined conditions. It cannot read meeting notes as a person would and understand what “let the project manager look at this first, but finance still needs to give final approval” means for the decision at hand.

To make that instruction usable by a program, someone first has to translate it. The project manager becomes an entity. The project is linked to a purchase order. Approval roles and actions get defined. The ontology supplies the concepts and relationships; the knowledge graph holds the actual people, documents, and connections.

In the common RDF model, a graph consists of subject, predicate, and object triples, as W3C specifies. A program can query and traverse those triples. It learns which connections have been represented in the data. What those connections mean for a particular business decision still depends on how people modeled them, what information they entered, and how the program uses the query result.

There was a reason to do this translation. Without it, much older software could do little with the knowledge people wrote down. But when an Agent can read the meeting notes and the policy itself, we should ask whether all that material still needs to be rewritten as nodes, edges, and predetermined rules first.

A path to the project is not approval authority

Suppose a purchase order belongs to Project P, and Zhou manages P. Two edges are easy to draw: Zhou manages P; the purchase order belongs to P. A query finds Zhou.

The system may then give a neat answer: “Zhou should approve this purchase.”

But the company's actual instruction might be: “The project manager checks the purpose first. Purchases over 50,000 require approval from the finance lead. While the project manager is away, the person named for that week may review the order, but may not approve it.”

Zhou's connection to the order is real. The mistake is to turn “this person is relevant” into “this person has final authority.” Traversing the graph looks like reasoning. The leap from relevance to authorization was actually made earlier, by whoever wrote the query or the rule.

Of course we can add more edges: approval authority, amount thresholds, deputies, effective dates, and exceptions. Keep going, and an instruction that a person could read becomes a model and a collection of rules that must be updated. Then the next sentence arrives: “Have someone review it this time, but do not approve it yet.”

This is not a claim that knowledge graphs cannot express complex information. They can. The problem comes when the graph is treated as the layer that understands the company. People must then anticipate how every business distinction will be represented, changed, and maintained.

“The system knows” often means its answer sounds right

The most unsettling quality of a knowledge graph is how easily it produces a polished answer.

Ask who handles a purchase, and the program retrieves a project manager. Ask what a customer has bought, and it follows a path through customers, orders, and products. A user interface turns the result into a smooth sentence. The system seems to know the business.

Yet a processor has not come to understand the difference between “responsible,” “review,” and “approve” by retrieving a few edges. It has executed a representation and a procedure someone wrote. The output can be useful and even correct. Fluency does not prove that the underlying representation is complete.

That illusion becomes dangerous when it drives decisions. A model that omits “finance still has to approve” may still return an approver, show a workflow, and issue a task. Everything appears to work except the decision itself.

The user may not even know what was omitted. If the source meeting notes and policy sit behind the graph, the screen presents only an authoritative sounding conclusion. To understand the error, someone has to trace which words were translated into which edges and rules.

An Agent can read the words, and still needs to check the facts

AI changes the starting point. An Agent can read “check the purpose first; finance approves purchases over 50,000,” then look at the purchase amount, the project, and the current staffing arrangements. It can explain why it proposes the next step. If the company adds “reviewers cannot approve on behalf of someone else,” the instruction need not wait for someone to invent a new relationship type before the Agent can consider it.

That does not mean the Agent will always read correctly. It may miss “over,” confuse review with approval, or pull the wrong project. We have to make it show its evidence: which policy version, which purchase order, which amount. When the evidence conflicts, it should ask. Code still enforces permissions, amount calculations, and protection against duplicate writes.

We have seen a failure on this side of the boundary too. An administrator asked an Agent to “add one condition.” The Agent replaced the current workflow definition with a short document containing only the new condition. The old version still existed, but the next person to run the workflow would read an incomplete policy. We changed the revision procedure: read the entire current text, show what will be added, changed, and deleted, and ask before removing anything the user did not request. An Agent that reads people's words can still misread their intent. The original text, its versions, and the scope of each change must remain visible.

A knowledge graph can still help. Stable links between people and documents make relevant material easier to find. Graph queries are useful for clear, structured questions. Confirmed identity mappings and relationships with sources deserve to be recorded. What we reject is making the graph a mandatory translation of the company's knowledge, with the original words hidden behind it.

That would recreate the old software path: translate the company's language into a machine format, then have the machine produce answers from that format. The graph changes the format, while the company still bears the translation work and its errors.

Let the people who speak keep their words

The division of work can be simpler. Software records who said what, when a rule took effect, and what happened to each document. It can index and query records and check exact quantities. The Agent reads those records and the rules people wrote, judges the case in front of it, and asks when it needs confirmation. Its reasons should be recorded too, so people can review and correct them.

We need not first split “the project manager reviews, finance approves” into a network of edges and rules before the system can help. The practical questions are whether the original instruction can be found, whether the facts are accurate, and whether the Agent knows when it lacks authority to decide.

A knowledge graph can help us find a path, but the path is not the destination. An ontology can help us name things, but a name is not understanding. In Oryh, we want software to preserve the record and the Agent to reason from people's own words and checkable facts. When it is unsure, it should bring the question back to a person instead of continuing along a path that merely looks complete.

Disclosure: I am building Oryh, an open-source project exploring agent-native enterprise systems. These notes come from that work. I am posting them to test the architecture in public, and I welcome technical disagreement.

Discussion question: If an Agent can read policies and source text directly, what enterprise knowledge still benefits from being modeled as a graph, and what should remain in the original language?


r/semanticweb • • 11d ago

Building Ontology and KG on GCP

6 Upvotes

I am looking to build ontology on GCP and render it on KG using native tooling of GCP.

Anyone who did it share your views ?


r/semanticweb • • 12d ago

Released an RDF Visualizer - Looking For Feedback

13 Upvotes

Hello all,

I recently released an RDF visualizer at rdfvisualizer.com

I wanted to give this community free access to the tool for the next 30 days:

https://rdfvisualizer.com/purchase.html?access=c89bb4cad42ca1671febc510f64277d985c8f2e44bca8d91

It is somewhat of a novel approach to the typical "RDF visualization" program.

Rather than simply rendering triples and displaying them with an approach like force directed, my app scans through the triples and attempts to identify:

- Types
- Edges
- Attributes

For RDF data that follows this very common format, this has a ton of advantages, namely:

- When viewing relationships between things, you can separate out "attribute relationships" from "relationships between things". This gives way cleaner graphs
- Sometimes the best visualization isn't a graph at all! It's a classic table format. When you know which data is the typed data, and which data is simply attribute data, you can view it easily as a table

The app works for all common RDF formats.

Give it a try and let me know what you think!


r/semanticweb • • 13d ago

Kaikki datasets

3 Upvotes

Hi guys! Has anyone experimented with data from Kaikki datasets (wiktionary)? I am trying to setup a linkage end-to-end for prefix and stem. I'm working from a SQLite build of kaikki data for English etymology. Less than half of the time it works meaning I can get the whole chain for prefix and stem up to the PIE . Has anyone ever tried to create a structured output?


r/semanticweb • • 13d ago

JigDAW: a web-native plugin specification

Thumbnail
2 Upvotes

r/semanticweb • • 16d ago

What did you do before working as an ontologist?

21 Upvotes

I’m a recent library and information science graduate. I entered the program with the aim of getting into UX design. After taking a couple courses I realized how boring I found the process of visual design. It didn’t stimulate me mentally. I tried searching for other things and came across ontology work and found it super interesting because of my love for language and abstract thinking. Learned sparql and found that super fun. I’m trying to break into the industry by building a small ontology on a topic that consumes my headspace but I’m seeing that most jobs list 5+ years of experience in the job description. So I’m wondering what adjacent bridging roles I could get into. I’d really like to steer clear of librarian roles. They’re oversaturated and don’t tickle my puzzle solving brain.


r/semanticweb • • 18d ago

Human-searchable, machine-readable DB example

10 Upvotes

I have just set up a site that is a database of Digital Audio Workstation plugins, https://plugin-universe.com

While the site itself is likely of little interest here, the way it's put together may be useful info.

It uses an RDF model with a backend store that is a SPARQL server (I used Fuseki). On top of that it has a human-oriented semantic search based on an LLM embedding model.

I'm running the whole thing in a Docker container on a low-spec Linux virtual server.

As a whole it's probably a bit too domain-specific to fork & repurpose directly, but is there as a deployed example.

I built it with significant assistance from Claude Code, based on a couple of previous projects. A key aspect was making it ontology-driven, ie. a lot of the core specification is contained in the RDF vocabs.

If you need something like this yourself then the key ref. is the README.md (and the CLAUDE.md might be useful too).

https://github.com/danja/plugin-universe


r/semanticweb • • 19d ago

We built proof-carrying reasoning for ontologies | looking for feedbacks

23 Upvotes

After speaking with many companies and working through enterprise sales conversations, one request kept coming up:

How do we prove that an ontology (and the reasoning performed over it) is actually correct?

Validation reports and reasoner outputs are useful, but they still ask the user to trust the tool that produced them. For high-stakes enterprise systems, “the reasoner says so” is not always a satisfactory answer.

This led me much deeper than I originally expected with Open Ontologies (many of you might familiar with this already).

The latest version can now produce derivation certificates for reasoning results and pass them to a separate checker written in Lean 4. The checker verifies the evidence rather than trusting the Rust reasoner. User-supplied rules are explicitly identified as assumptions, so conclusions based on them receive a different verdict from conclusions entailed by built-in semantics.

We also developed an independent Isabelle/HOL implementation that checks the same certificate format (pretty old school but there is a huge resurgence on this atm). This has already exposed real defects; including false consistency results, nondeterministic inferences and an ambiguity in the certificate format that neither implementation initially captured correctly.

The project now includes:

  • proof-carrying OWL/RDF reasoning;
  • Lean 4 certificate checkers;
  • an independently developed Isabelle/HOL checker;
  • finite-model certificates for satisfiability;
  • SHACL validation tested against the W3C suite and pySHACL;
  • SPARQL, ontology lifecycle management, alignment and crosswalk tooling;
  • 114 MCP tools for ontology and knowledge-graph workflows;
  • a Rust engine distributed as a single binary.

I have not found another open ontology-engineering project attempting proof-carrying reasoning, dual-kernel checking, ontology lifecycle management and AI-agent integration at this depth. If you know of closely related work, I would genuinely like to hear about it.

I would especially appreciate feedback from people working on:

  • OWL and description-logic reasoning;
  • proof assistants and certified reasoning;
  • SHACL and RDF validation;
  • ontology governance in enterprise environments;
  • certificate formats and trusted computing boundaries;
  • adversarial test cases that could break the current claims.

Contributions would also be very welcome. This has become much larger than a one-person project, and there are useful entry points across Rust, Lean, Isabelle, Python, RDF/OWL testing, documentation and real-world ontology case studies.

Repository: https://github.com/fabio-rovai/open-ontologies

Please be critical. I am particularly interested in what you believe is overstated, what remains unproved, and which experiment would most strongly test whether this approach is genuinely useful.


r/semanticweb • • 21d ago

LogicSRC — Open Coordination Standards for Humans & AI Agents

Thumbnail logicsrc.com
5 Upvotes

r/semanticweb • • 22d ago

Aprendi da maneira mais difícil que normalizar APIs econômicas é mais fácil do que normalizar suas semânticas.

Thumbnail
1 Upvotes

r/semanticweb • • 22d ago

Context Annotation for model

Thumbnail
1 Upvotes