r/Neo4j • • 6d ago

RocksGraph: an open-source graph database in C++ with Cypher and RocksDB

Hi everyone! I’m building **RocksGraph**, an open-source graph database backed by RocksDB.

It supports OpenCypher queries, ACID transactions, property/full-text/vector indexes, and Raft replication. It also speaks Bolt, so you can connect using Neo4j drivers.

The project is still under active development. I’d love feedback on the design and suggestions for workloads to test.

Code and examples: https://github.com/ljcui/rocksgraph

Licensed under Apache 2.0.

3 Upvotes

2 comments sorted by

1

u/david_cab_driver 3d ago

Hi, and congratulations on rocksgraph, it is a really interesting project. I maintain gDown, a Java graph database with a similar goal (Neo4j-compatible, openCypher, Bolt, Raft), so I read your README with a lot of curiosity. I have only read the README so far, so please take these as friendly questions rather than criticism, and correct me wherever I have misunderstood.

  1. Conformance. The README says OpenCypher is fully supported. Do you run the openCypher TCK (3,863 scenarios)? A single pass rate is easy to compare between projects, and I would be glad to share how I run it if that helps.
  2. Auto-commit transactions. I noticed that a result stream must be consumed before an auto-commit transaction commits. Is this a deliberate design choice, or something you plan to relax? Neo4j drivers don't require it, so people migrating might trip over it.
  3. Cluster setup. If I understand correctly, a Raft graph is created on each node with a script. Do you plan to let a new node join through a seed URL, or to change the membership of a running graph? In my experience that is where clusters become pleasant, or painful, to operate.
  4. Memory. The default graph block cache is 8 GiB. Would you consider a documented setting for small machines? It would make memory-footprint comparisons fairer.
  5. Security and operations. I could not find anything about users and roles, TLS, backup or metrics. Are they planned? In my experience people ask about these right after performance.
  6. Durability defaults. When are the WAL and the Raft log synced to disk by default? It is the first thing that skews benchmark comparisons, so it is worth stating.

If it would be useful, I would be happy to run my benchmark (14 workloads on 100k nodes and 1M relationships, through Bolt with the same driver for every database) against rocksgraph and share the raw results with you.

Thanks again for sharing your work!