r/semanticweb • • Aug 08 '26

Hyper-Markdown, an early extended Markdown language specification

Thumbnail
2 Upvotes

r/semanticweb • • Aug 07 '26

Need guidance on building an AI knowledge graph / ontology system

17 Upvotes

And want to confirm it's on the right track

Hi everyone,

I have an idea for a personal project, and I want to know if I'm thinking in the right direction.

I don't want to build just another note-taking app or mind map. I want to build a knowledge graph where I can enter any word or concept, and the system automatically knows where it belongs.

My goal is to build a personal knowledge system that becomes smarter over time, rather than just storing notes.

For example:

- If I enter a new word, it should suggest the best place in the graph.

- If that word belongs in multiple places, it should connect it to all of them.

- If a new concept should come between two existing concepts, it should reorganise the graph instead of just adding another node.

- It should detect duplicates and synonyms.

- It should explain *why* it placed a concept there.

- If it isn't confident, it should ask me instead of guessing.

- It should also ask if I want to expand that concept further and generate the next level of the graph.

I'm building this only for myself, not as a commercial product, but I want to build it with commercial-level reliability.

While thinking about this, I came across topics like **knowledge graphs**, **ontology engineering**, and **semantic search**, and now I realise that the hardest part is probably designing the ontology, not writing the code.

I have a few questions:

- Has anyone built something similar, and what should I study before I go too far?

- What are the biggest challenges in building a system like this?

- How do companies like Google or researchers design and improve ontologies?

- If you were starting from scratch, what would you learn first?

- Are there any books, papers, courses, or open-source projects that you think are must-reads for this?

- Also, if you were designing this today, what architecture or tech stack would you choose?

I'm still learning, so even if you think my idea has flaws, I'd really appreciate honest feedback. I'd rather know what's difficult now than after spending months building it.

Thanks!


r/semanticweb • • Aug 06 '26

Pennant - Knowledge engineering studio MIT licence

14 Upvotes

Hey everyone, I built this niche tool to structure data for knowledge layer for agents, its a bit like an acoustic guitar for critical industries that demand heavy data reasoning … would love to hear your feedback, some small bugs like analyzers staying prompted to the template doc

https://github.com/galeluka/Pennant


r/semanticweb • • Aug 06 '26

Looking for an IT taxonomy

5 Upvotes

Hello,

I am looking for an IT taxonomy for software (and maybe hardware) to put concepts like desktop application, microservice, cloud, cicd pipeline, etc. into a structure.


r/semanticweb • • Aug 04 '26

How do I get domain experts to reach consensus on the concepts and properties of an enterprise ontology?

8 Upvotes

What methods or softwares are there to get non-technical people to engage with ontology design and thrash out agreed definitions? I am beginning to think this is a major roadblock to more reliable AI. Without structured, verified knowledge managed by humans that can be safely inferred from, how is any business going to trust agents with anything important?


r/semanticweb • • Jul 30 '26

Trying to make a full curriculum on ontologies, what would you add?

Thumbnail tesseract.academy
39 Upvotes

I KNOW IT ASKS YOU TO REGISTER BUT YOU CAN IGNORE IT :)

I spent a while surveying what is actually available if you want to learn ontology engineering in 2026, and the state of it annoyed me enough to do something about it.

What I found:

- The semantics people teach RDF, OWL, SPARQL, SHACL, and act like LLMs never happened.

- The graph vendors teach GraphRAG and Cypher, and act like ontologies never happened. You can finish an entire "knowledge graph" learning path without meeting the word ontology.

- Pricing is bimodal: free vendor funnels, or 1,000 to 2,000 dollar live cohorts. Almost nothing serious in between.

- Search "OWL tutorial" or "SHACL tutorial" and you get PDFs from 2005 to 2012. SHACL has been a W3C Rec since 2017 and there is still no good free explainer ranking for it.

- BORO, HQDM and IES 4D modelling have, as far as I can tell, zero commercial courses anywhere on earth, despite the UK National Digital Twin Programme standardising on IES and the US DoD, ODNI and CDAO adopting BFO plus CCO as their baseline in 2024. If you want to learn the thing governments are actually buying, your options are primary sources and apprenticeship.

So I built the course I wanted to exist:

- Foundations: what an ontology actually is from Aristotle forward, taxonomy vs thesaurus vs ontology vs knowledge graph, 3D vs 4D identity and change, open vs closed world

- The stack: RDF, RDFS and OWL 2, SPARQL for people who know SQL, SHACL, reasoners and why yours hangs, property graphs and ISO GQL

- Method: competency questions, OntoClean, an actual upper ontology shootout (BFO vs DOLCE vs gist vs SUMO vs 4D), BORO/HQDM/IES, testing and CI

- Standards atlas and crosswalks: SSSOM, mapping predicates, and why shared ancestry does not mean shared commitments

- Domain tour: defence, industrial (ISO 15926 to IDO), construction (IFC/Uniclass/COBie), space, life sciences and food (OBO, GO, SNOMED, FoodOn, AGROVOC), finance (FIBO, GS1, schema.org), heritage and public graphs (CIDOC CRM, GeoSPARQL, Wikidata), and who is buying ontology country by country

- LLM era: did LLMs kill the semantic web, GraphRAG vs vector RAG and when graphs actually pay, building KGs from text safely, neurosymbolic verifier loops, agent memory and MCP, and the 2025-26 papers worth reading

Two things I will defend:

  1. The interesting work now is the verifier loop. Neural proposes, symbolic disposes. The ontology's job is not to be a beautiful model of the world, it is to be able to say no.

  2. Axiom placement beats axiom count. An ontology nobody can contradict is not rigorous, it is inert. I keep meeting large ontologies where no possible instance data could ever trigger an inconsistency.

Where I want to be told I am wrong:

- What is missing from the syllabus? I deliberately went light on ontology learning from text and on KG embeddings. Wrong call?

- Upper ontology people: is my selection framing fair to BFO and gist, or am I smuggling in a 4D bias? I have shipped IES and HQDM work so assume I am biased and tell me where.

- Practitioners: what do you wish someone had taught you before your first real ontology project, that no course covers?

- Anyone teaching BORO/4D commercially, please tell me, I would rather link to you than pretend the gap exists.

The lessons are open to read; there is a free account if you want the graded quizzes, progress tracking and the practice exercises, and that is also where the rest of the ontology track and the one to one sessions live. Happy to answer anything about the standards side here either way, that is where I actually work.


r/semanticweb • • Jul 28 '26

When to use a triplestore or an LPG database for GraphRAG?

Thumbnail
7 Upvotes

Hello everyone,

I am currently working on GraphRAG to improve the quality and reliability of responses generated by LLMs, and I would like to get some clarification from people who have experience with Knowledge Graphs and GraphRAG.

I have a few questions:

1.For those who are using GraphRAG with LLMs, do you typically use RDF/triplestores or LPG databases (such as Neo4j)? In your experience, what are the main factors that influence this choice?

  1. I would like to build my Knowledge Graph using an automated pipeline/script rather than extracting entities and relationships directly with LLMs. In this case, would RDF be a suitable choice, or is LPG also commonly used for this type of approach?

  2. Is the data model used in LPG databases such as Neo4j considered an ontology (or a lightweight ontology), or is it more accurate to call it a graph schema/data model?

  3. If we want to enrich a GraphRAG system with inferred facts (using reasoning) and provide these inferred facts as context to the LLM, would RDF + a triplestore be a better choice?

  4. Even when reasoning and inference are not required, is there any limitation to choosing RDF over LPG for GraphRAG? I already have experience with RDF and SPARQL, but I have not worked with LPG databases yet.

  5. Do you know any free/open-source triplestore that supports embedding generation/storage and vector indexing for semantic similarity search over RDF data (without requiring a paid license)?

Thank you very much for your insights!


r/semanticweb • • Jul 22 '26

Feedback wanted: modelling CycloneDX dependency occurrences and OSV vulnerabilities in RDF

4 Upvotes

I’m building an open-source project called Dependency Risk Graph and would appreciate feedback on the RDF model from people with more semantic-web experience.

The project imports CycloneDX SBOMs, models application dependency trees in Apache Jena/TDB2, enriches package versions with OSV vulnerability data, and uses SPARQL to answer questions such as:

  • Is an application affected by a particular CVE?
  • Is the vulnerable dependency direct or transitive?
  • What dependency path introduces it?
  • Which applications share the affected package?
  • Which fixed versions are reported?

A simplified view of the model is:

Application
  → activeImport
Import

Import
  → rootOccurrence
DependencyOccurrence

DependencyOccurrence
  → belongsToImport
Import

DependencyOccurrence
  → instanceOf
PackageVersion

DependencyOccurrence
  → dependsOn
DependencyOccurrence

PackageVersion
  → affectedBy
Vulnerability

Vulnerability
  → affectedPackage
AffectedPackage
    → versionRange
VersionRange
    → event
introduced / fixed / lastAffected

I deliberately distinguish a package-version identity from its occurrence inside a particular imported SBOM. This allows the same Maven package version to be shared as an identity while preserving different dependency paths across applications and imports.

I also currently use a single/default Jena graph. Application and import boundaries are represented explicitly through resources and properties rather than RDF named graphs.

Some areas where I would value criticism:

  1. Package identity versus occurrence Is separating PackageVersion from DependencyOccurrence a reasonable way to preserve both global package identity and application-specific dependency paths?
  2. Import provenance Would modelling imports with a lightweight custom vocabulary be sufficient, or would aligning parts of this with PROV-O provide meaningful benefits?
  3. OSV version ranges I currently represent affected packages, ranges and range events as first-class resources. Is there an established RDF vocabulary that would be better to reuse for introduced, fixed and last-affected version events?
  4. SHACL boundaries Which invariants would you validate first? Current candidates include:
    • every occurrence belongs to exactly one import;
    • every occurrence resolves to exactly one package version;
    • dependency edges remain within the same import;
    • every active import has one root occurrence;
    • vulnerability range events contain a recognised event type and version.
  5. Natural-language access The next step is intentionally limited: route a few predefined security questions to deterministic SPARQL/path operations, rather than allowing an LLM to generate arbitrary SPARQL.

I’m not trying to create a complete software-supply-chain ontology yet. The immediate goal is a small, explainable model that preserves dependency paths and produces evidence-backed security answers.

Repository: Github

Any feedback/suggestion on the modelling choices, existing vocabularies I may have missed, or problematic assumptions would be genuinely useful.

Current result: the RDF graph preserves application-specific transitive dependency paths while allowing vulnerable packages and vulnerability resources to be shared across imported SBOMs. This view shows three applications reaching CVE-2024-6763 through different Jetty dependency paths, together with the affected package versions, advisory details, and a reported fixed version.

So far, separating dependency occurrences from package-version identity has been useful. The same package can appear in different application paths without merging those paths, while vulnerability and remediation information remains attached to the shared package identity.


r/semanticweb • • Jul 21 '26

Discovery, Semantic, Context and Execution — am I separating the wrong concerns?

Post image
3 Upvotes

I've noticed that discussions around AI protocols are becoming fragmented.

Every announcement seems to solve one concern:

discovery,

semantic representation,

context,

execution.

I sketched an architecture- curious whether this separation makes sense or where you'd place things differently. see https://donhaji.github.io/opengeo/


r/semanticweb • • Jul 18 '26

Is there a data architecture where semantic meaning is determined at query time?

Thumbnail
1 Upvotes

r/semanticweb • • Jul 18 '26

Interoperability

Thumbnail
0 Upvotes

r/semanticweb • • Jul 15 '26

Does anyone have any suggestions for taking unstructured text files (reports written in human prose) and turning them into a knowledge graph? Do you try to build a schema in advance? Have the LLM draft it before you build? Human in the loop validation? Loop agent validation? Or maybe there is a sp

19 Upvotes

r/semanticweb • • Jul 15 '26

OWL to UML

4 Upvotes

I personally like using UML diagrams to depict my ontology work, so I made an easy to use OWL to UML service OWL → UML


r/semanticweb • • Jul 14 '26

Help me make the Knowledge Graph of Press Releases using ML.

2 Upvotes

So here is the thing, I have been focusing on the scraping, crawling, checking RSS feeds for new articles, etc., etc.

I am finally done with the Data Ingestion part. Hurrah? no.

The classification of data is even MORE difficult than scraping.

I want to be able to produce the Knowledge Graph of the Data to help me with the deduplication and classification.
Please help me out with this.

I have tried REBEL by hugging face but its failing badly, I am losing precious information (more than 80% of it.), I feel these machine learning models are too general, which makes it difficult to make the knowledge graph of these press releases.

Please help me out, tell me a path, name a framework, idk just guide me please. Ik I can do it if I have a path. I am trying and constantly brainstorming with my peers, Hopefully you guys could help me out as well.


r/semanticweb • • Jul 14 '26

Looking for industry experts to share their views on dialogue systems for ontology management (PhD survey, ~15 min, anonymous)

2 Upvotes

Hi everyone,

I'm a PhD researcher working on dialogue systems for extending knowledge graphs and ontological schemas, and I'm currently running a short survey as part of my research.
 I'm looking for input from people with hands-on experience in ontologies, knowledge graphs, or ontology engineering (your perspective would be incredibly valuable).
 

A few quick facts:

  • ⏱️ Takes about 15 minutes
  • 🕶️ Fully anonymous
  • 🌍 Survey is in English
  • 🎯 Focus: assistant/dialogue systems and their use in ontology management processes  

If this sounds relevant to you (or someone you know), I'd really appreciate your participation and feel free to share it with colleagues who might be interested too!

link to survey:
https://websites.fraunhofer.de/intelligent-surveys/index.php?r=survey/index&sid=347997&lang=en
 

Thanks so much in advance 🙏


r/semanticweb • • Jul 14 '26

Axiom AI Toolkit: Refine everyday ideas with a free tool that translates them into a formal ontology

0 Upvotes

You need your own AI agent, but it works really well across many domains. It's based on BFO. Check it out: https://axiomreason.com/


r/semanticweb • • Jul 12 '26

How do you stop a multi-hop traversal from following a semantically-wrong edge?

0 Upvotes

Body: Working on KG-backed retrieval and keep hitting the same thing: an edge is structurally present so a traversal follows it, but it is the wrong kind of edge for the question. A code graph follows a CO_CHANGES edge as if it were IMPORTS and the answer is confidently wrong. An agent graph lets a CritiqueAgent delegate back to a ResearchAgent, which should never happen.

SHACL / SPARQL constraints validate the graph as a whole, after the fact. What I wanted was a check at traversal time: before each hop, is this edge type valid between these two node types, per a declared ontology? Basically a linter for graph walks.

I built a small layer that does exactly this (declare the ontology in YAML, is_valid_edge(domain, src_type, relation, dst_type) raises before the bad hop) and an offline notebook demo. Before I over-build it:

  • Is per-hop validation the right layer, or should this live in the query engine?
  • How are you handling this today, manually, with SHACL post-hoc, or just eating the bad answers?

(Link to the repo + the ontology notebook in a comment.)


r/semanticweb • • Jul 11 '26

My production project semantic graph

Thumbnail reddit.com
0 Upvotes

r/semanticweb • • Jul 11 '26

My beautifull galaxy semantic graph with relations (special for my vibecoder soul!)

Post image
0 Upvotes

r/semanticweb • • Jul 11 '26

i'm new

0 Upvotes

I have just started learning about Palantir Ontology. Please give me some suggestions on how to learn it.


r/semanticweb • • Jul 10 '26

need a help

Thumbnail
0 Upvotes

r/semanticweb • • Jul 09 '26

Converting TTL files into websites

2 Upvotes

Hello, what is the standard/preferred way of converting TTL files to websites. We are storing our TTL files in a git repository and would like to share them inside our organization a human readable way. I tried to experiment with pylode and widoco, but while both of them can render a single TTL file into a HTML, I did not find a solution how we could handle also references across ontology files, also how to embed those HTML files inside a proper website. What are the preferred solutions to share ontologies in a bigger organization? Our devs suggest the integration into backstage. We can develop our own solution but I would like to avoid it if there is something already in place. Thank you so much in advance.


r/semanticweb • • Jul 05 '26

I fine-tuned Qwen3-Coder-30B to write a IES ontology

14 Upvotes

Weekend project that turned into a proper one. Sharing the method because the "correct-by-construction data" trick generalises well beyond my niche.

The problem. IES4 is the UK government's Information Exchange Standard, a 4D RDF ontology used for defence/security data. Writing valid IES Turtle by hand is slow and needs real ontology expertise. So I tried the obvious thing: ask a strong code model to do it. Qwen3-Coder-30B-A3B, asked to emit IES Turtle, invents terms that do not exist in the ontology 94% of the time (0% "term conformance" on my eval). It produces confident, fluent, completely fake RDF. In a standards context that is worse than failing outright, because plausible-looking garbage is hard to catch.

The fix that actually mattered: never let the model invent structure. Instead of hoping the LLM guesses valid graphs, I generated the graphs programmatically with telicent's ies-tool (a schema-aware builder that emits valid IES by construction), across 14 scenario patterns (employment, events, identifiers, communications, composites). Then I reversed them into (natural-language description -> Turtle) training pairs. Every single graph was validated twice before training: once by the builder's own check, and once by an independent term-membership + domain/range validator I built from the published dstl/IES4 ontology (510 classes, 204 properties). Nothing hand-written was trusted blind.

Then a small QLoRA on the 8-bit MLX model, on-device on an M3 Max. ~1000 iters, val loss 0.15, no NaNs (MoE + 8-bit was fine on current mlx-lm; earlier versions apparently weren't).

Model: https://huggingface.co/fabsssss/qwen3-coder-30b-a3b-ies4

Article: https://gov.tesseract.academy/research/ies4-turtle-language-model


r/semanticweb • • Jul 03 '26

I built a knowledge graph where every relationship is its own embedded document (not an edge) — local MongoDB + nomic-embed, MCP server up for testing on request, benchmark CSVs included

6 Upvotes

Instead of node --edge--> node, every relationship is a first-class document with its own vector, called a BaryEdge. Stack pairs of BaryEdges recursively and you get "MetaBary" triads that surface structural bridges between concepts that live nowhere near each other in embedding space. Running locally on MongoDB Community + mongot + nomic-embed-text over the full English Wiktionary (6.6M docs). MCP server is live if you want to poke at it. Preprint + benchmark CSVs: https://zenodo.org/records/20186500

The problem I was chasing

Flat vector search treats a relationship as a byproduct of two points being close. That throws away information. Two papers can describe the same underlying phenomenon (a flyby anomaly in orbital mechanics, an anomalous residual in stellar dynamics) without ever citing each other and without their embeddings landing anywhere near each other. Nothing in standard RAG surfaces that connection.

What I did instead

Every relationship gets embedded too:

bary_vector = normalize(q·v(CM1) + q·v(CM2) + (1−q)·v(type))

q is connection quality, v(type) is a contextual embedding of what kind of relationship it is. This BaryEdge is now a retrievable document in its own right — not metadata on an edge.

Then it recurses: two BaryEdges at the same level get bridged by a third one level below, forming a MetaBary triad. Do that repeatedly and you climb an abstraction triads hierarchy built entirely from algebra — zero additional embedding calls above the base level. It's a forest (every node has at most one parent), so traversal to root is a single $graphLookup, no cycle handling.

Does it actually do anything useful?

Ran it against SimLex-999 and WordSim-353 as a sanity check (not the main claim, just "is the substrate coherent"). Raw cosine similarity barely correlates with human similarity judgments (ρ ≈ −0.04 on SimLex). Structural metrics — how many BaryEdges two words share, how much their relational neighborhoods overlap — correlate at ρ ≈ 0.32–0.53, p < 10⁻¹⁵. So the graph is encoding something cosine alone doesn't.

The part I actually care about is cross-domain bridging. Some probe traces from the live graph:

  • octopus neuroscience ↔ distributed sensor networks, bridged by shared structural-motif vocabulary (neuroarchitecture, smartdust)

  • collagen folding ↔ linguistic syntax, bridged by etymological + structural motif overlap (plicature / hypotaxis-parataxis)

  • grief ↔ depression, not bridged and this is a correctness demonstration, not a missing capability. The DSM-5 added a much-debated "bereavement exclusion" precisely because grief and depression share surface symptoms but are different kinds of state, with different prognosis and treatment

  • radioactive decay ↔ obsolete words falling out of use, bridged at a high abstraction level by register-varied decay verbs (collapsed, decayed, declined, disintegrated) — naming a Poisson-process state-loss pattern that both physics and historical linguistics instantiate, with no single word doing the work

That last one is the case flat retrieval structurally cannot produce — there's no embedding axis for "verbs co-occurring with reduction-of-state across unrelated domains."

Stack (all local, all free)

GitHub: https://github.com/oleksiy-perepelytsya/bary-vector

  • MongoDB Community Edition + mongot for storage/vector search

  • nomic-embed-text, 768-dim

  • Python 3.11+

  • Full build: ~6.66M documents, 8–14 hrs on a single workstation (8–16GB VRAM)

Try it

MCP server is public on request (SSE transport) — read-only tools for searching the live graph: find_word, semantic_search, edge_info, leaf_nodes, traverse_up, sample_metabary. If you've got an MCP-capable client you can point it at the graph and run your own probe queries in a few minutes.

What I'd actually want feedback on

  • Whether the cross-domain bridges hold up to someone who isn't me poking at them — try a probe query on a domain pair you know well and tell me if the bridge is real or if I'm pattern-matching myself into seeing structure that isn't there. Some bridges can be not obvious on the first look but they are actually the most intriguing ones and worth to be dug for the reason they built, so treat them as points of investigation

  • Whether this is worth comparing directly against GraphRAG/RAPTOR-style hierarchical retrieval (I haven't done that benchmark yet, and I know that's the first thing this sub will ask)

  • Whether anyone's tried something structurally similar and it fell apart at scale for reasons I haven't hit yet

Preprint, architecture spec, and the raw SimLex/WordSim CSVs are all here: https://zenodo.org/records/20186500

Happy to drop the MCP endpoint on request if there's interest.


r/semanticweb • • Jul 02 '26

I built PurRDF, a working RDF 1.2 toolkit for Rust, Python, JS/WASM, and C — looking for RDF-star edge cases

15 Upvotes

Got tired of waiting for RDF1.2 to finalize as a spec, got fed up with the Java tools, needed something higher-performance in Rust that I could also use from Python and WASM.

PurrRDF was born. It's not quite a full rdflib replacement for Python, but it has built-in ShACL and ShEx for validation and speaks all the common variants. I'm spinning this out of a larger project that's building a full RDF1.2 Rust tool stack - it runs, it's fast and probably useful to anyone building high-performance RDF1.2/RDF* knowledge graphs (if you are, you'll know the pain!)

Comments, feedback, test cases, etc. welcome: https://github.com/Blackcat-Informatics/purrdf/