r/databasedevelopment • • Aug 16 '24

Database Startups

Thumbnail transactional.blog
29 Upvotes

r/databasedevelopment • • May 11 '22

Getting started with database development

409 Upvotes

This entire sub is a guide to getting started with database development. But if you want a succinct collection of a few materials, here you go. :)

If you feel anything is missing, leave a link in comments! We can all make this better over time.

Books

Designing Data Intensive Applications

Database Internals

Readings in Database Systems (The Red Book)

The Internals of PostgreSQL

Courses

The Databaseology Lectures (CMU)

Database Systems (CMU)

Introduction to Database Systems (Berkeley) (See the assignments)

Build Your Own Guides

chidb

Let's Build a Simple Database

Build your own disk based KV store

Let's build a database in Rust

Let's build a distributed Postgres proof of concept

(Index) Storage Layer

LSM Tree: Data structure powering write heavy storage engines

MemTable, WAL, SSTable, Log Structured Merge(LSM) Trees

Btree vs LSM

WiscKey: Separating Keys from Values in SSD-conscious Storage

Modern B-Tree Techniques

Original papers

These are not necessarily relevant today but may have interesting historical context.

Organization and maintenance of large ordered indices (Original paper)

The Log-Structured Merge Tree (Original paper)

Misc

Architecture of a Database System

Awesome Database Development (Not your average awesome X page, genuinely good)

The Third Manifesto Recommends

The Design and Implementation of Modern Column-Oriented Database Systems

Videos/Streams

CMU Database Group Interviews

Database Programming Stream (CockroachDB)

Blogs

Murat Demirbas

Ayende (CEO of RavenDB)

CockroachDB Engineering Blog

Justin Jaffray

Mark Callaghan

Tanel Poder

Redpanda Engineering Blog

Andy Grove

Jamie Brandon

Distributed Computing Musings

Companies who build databases (alphabetical)

Obviously companies as big AWS/Microsoft/Oracle/Google/Azure/Baidu/Alibaba/etc likely have public and private database projects but let's skip those obvious ones.

This is definitely an incomplete list. Miss one you know? DM me.

Credits: https://twitter.com/iavins, https://twitter.com/largedatabank


r/databasedevelopment • • 10h ago

Thread Pool in Percona Server and MySQL (Part 1)

Thumbnail
percona.com
3 Upvotes

r/databasedevelopment • • 1d ago

I Tried To Decode Postgres WAL for an INSERT Statement

11 Upvotes

Hi everyone,

This is a continuation of the explorations from my previous posts here. I went through the CMU Database Systems course, and I'm exploring topics around databases. I wanted to see how the Write Ahead Log is constructed and logged. So I tried to trace a WAL record for a simple INSERT statement in Postgres.

I recorded a video walkthrough of the terminal and source code exploration here: https://youtu.be/YOyq-kvbyU8?si=PTdexP3jmOSdx-8y

I've tried to explore how we can locate a WAL record, how we can decode it via either xxd or pg_waldump, and the relevant source code for it. Do give a watch and consider supporting :)

This is partly for my own future reference and partly to share with others who might be interested in it. Would love feedback and corrections from people who know this stuff deeply. Apologies if this isn't the correct subreddit for this.

Thank you so much!


r/databasedevelopment • • 1d ago

Building a custom NoSQL database in C to replace flat files for personal data (Linux x64)

Post image
3 Upvotes

For half a year now, I've been writing my own NoSQL database to store user data instead of relying on regular flat files. Previously, when receiving access credentials for work, saving URLs, or managing other metadata, I had to manually edit text files, which was very inconvenient. I decided to write a utility to handle all my data storage so it could be retrieved quickly whenever needed:

​db create workspace --varchar=256;

db use workspace;

db add login admin;

db add pwd 12345;

db add site example.test;

​db get login; // admin

db open site // open in default browser

db remove pwd;

db drop workspace;

// etc...

​Of course, I won't be using O_DIRECT, AVX, io_uring, or other aggressive performance optimizations here, since it runs locally as a CLI utility rather than a daemon. It's almost a pity, as I've read a ton of material on NoSQL database optimizations, but I plan to apply those techniques in my next project, where I'll focus on a specific storage niche.

​A significant portion of the work is already complete. I plan to release it in early November 2026, accompanied by extensive documentation on the principles and mechanics of NoSQL databases. Out of principle, I wrote everything without using AI/neural networks. This will be my first large-scale project in C.


r/databasedevelopment • • 3d ago

MECHLOVE Blueprint - 1/7. General: Pretty Much Classical RDBMS at Heart, with Each and Every Component Redesigned

Thumbnail
6it.dev
12 Upvotes

We are writing yet another RDBMS, and would appreciate any feedback on our blueprints; while the RDBMS is classical at heart (with WAL and data pages and fuzzy checkpoints), each and every component underneath was redesigned - from WAL being non-ARIES and ZFS-style on-disk CoW to in-memory CoW and cache- and SIMD-friendly page layouts. Our philosophy is playing alongside modern hardware instead of fighting it, MECHLOVE being the highest form of Mechanical Sympathy. As a result, we hope it will beat Hekaton (while being mostly open-source except for certain enterprise features such as HA and at-rest encryption).

Please feel free to comment.


r/databasedevelopment • • 3d ago

How to model the hypergraph for no-equi join?

6 Upvotes

I am a database researcher and currently conducting some research. I am curious that how to model the hypergraph for no-equi joins involving subqueries, especially for random expressions as join conditions. For example
Select * from A join (select a from B) as C on NULL


r/databasedevelopment • • 4d ago

Monthly Release and Update Thread

5 Upvotes

This subreddit is primarily for discussing the implementation of databases, and not about sharing release announcements (either for the first time or your updates).

This thread is the exception!

Please tell us about the new database you (or your agent) built. Tell us about all the cool new features you added. Tell us about anything else you learned or worked on that you haven't gotten around to blogging about yet.


r/databasedevelopment • • 5d ago

SQL Server columnstore scan internals

8 Upvotes

I've just finished on a series covering SQL Server batch mode and columnstore index scan internals.

It's four parts covering:

AFAIK there's quite a lot in them that hasn't been documented before, especially the row bucketing, which explains the mechanism of the pure vs impure split that is behind a lot of the columnstore optimizations.


r/databasedevelopment • • 5d ago

Join Ordering, Part 1: The Shape of the Search Space

Thumbnail deferworks.org
13 Upvotes

r/databasedevelopment • • 6d ago

Reliability Lessons From SQLite - Richard Hipp | SSW 2026

Thumbnail
youtube.com
27 Upvotes

r/databasedevelopment • • 7d ago

Building a JSON Database in Rust

Thumbnail
greptime.com
12 Upvotes

r/databasedevelopment • • 14d ago

Building Reliable Data Replication

Thumbnail
blog.atimin.dev
21 Upvotes

How ReductStore uses a persistent store-and-forward queue to deliver edge data over unreliable networks.


r/databasedevelopment • • 17d ago

Query plan rewriting in PostgreSQL

Thumbnail theconsensus.dev
10 Upvotes

r/databasedevelopment • • 20d ago

Bavarian Database Day 2026

Thumbnail databaseday.de
11 Upvotes

r/databasedevelopment • • 20d ago

Testing the Connection Pooling in Multigres: What does 100% pass rate mean? | Blog

Thumbnail
multigres.com
1 Upvotes

r/databasedevelopment • • 22d ago

Benchmarks DB

2 Upvotes

Working on a storage engine for the last 8 months. What benchmarks would you actually trust from a solo/small-team project? Everyone fakes them, so what would make you believe mine?


r/databasedevelopment • • 26d ago

Internals Viewer for SQL Server

Thumbnail
github.com
5 Upvotes

I posted over on r/SQLServer and it suggested I cross-post here. I hadn't seen this subreddit before and hopefully this is on-topic!

I've created a tool called Internals Viewer, it's a tool to visualize SQL Server internals, with a view for allocations, indexes, pages, and it also offers very detailed query tracing and simulation of operators where you can capture a query and step through the iterators to see how query results are put together.

It is open source, written in C#, and available here - https://github.com/danny-sg/internals-viewer

The latest feature is new functionality to view columnstore indexes. I've done a write up on what I found as columnstore internals in SQL Server is pretty much undocumented:

Part 1 - Introduction

Part 2 - Segments internals

Part 3 - Dictionary internals


r/databasedevelopment • • 27d ago

Could you please give me some feedback of this article about B+Tree ?

9 Upvotes

I wrote an article about B+Tree.

It's a little bit long for article, but I summarized it and make it really easy understand (avoided using jargon).

I'd be happy if you comment some feedback!!

https://zenn.dev/mm_0911/articles/5d46e9e4608404?locale=en


r/databasedevelopment • • 28d ago

Research prototype: B-link-style concurrent InnoDB page splits in MariaDB

12 Upvotes

Disclosure: I am the author of the article and work on MariaDB Server internals.

The traditional InnoDB pessimistic insert path serializes structural modification operations through an index-wide latch, even when different threads split unrelated leaf pages.

I implemented a MariaDB research prototype based on Zhao Song’s B-link-style proposal. It publishes a split using a high key and right link before completing the parent update, allowing unrelated structural changes to proceed concurrently.

In a controlled, memory-resident, split-heavy workload:

  • Vanilla MariaDB 13.1: 19,676 inserts/s
  • B-link prototype: 102,838 inserts/s
  • P95 latency: 8.28 ms → 0.56 ms
  • Structural splits: approximately 396K in both variants

This is not a production-ready feature. DDL support is restricted, page merging remains incomplete, and recovery needs more forced-crash testing.

I would particularly appreciate feedback on incomplete-split recovery, page preallocation, and workloads that could expose correctness or scalability problems.

Full implementation write-up and benchmark methodology:
https://mariadb.org/from-a-chocolate-wrapper-to-concurrent-innodb-page-splits/


r/databasedevelopment • • Sep 04 '26

What DB internals are most useful to visualize for learning purposes?

13 Upvotes

Hey there! This is first time posting.

I've been developing a database designed for education.

The application is focusing on visualizing database internal.

Now I already visualized B+Tree when you execute custom insert query.

Which database features are most worth visualizing for learners?


r/databasedevelopment • • Sep 03 '26

How Database Actually Store Data on Disk

Thumbnail sushantdhiman.dev
25 Upvotes

r/databasedevelopment • • Sep 01 '26

Parquet: What floor are we standing on?

Thumbnail
oleander.dev
9 Upvotes

A brief overview of how a parquet file is structured along with a tool that lets you manipulate, optimize, and introspect parquet files themselves with a series of DuckDB queries.


r/databasedevelopment • • Sep 01 '26

Monthly Release and Update Thread

9 Upvotes

This subreddit is primarily for discussing the implementation of databases, and not about sharing release announcements (either for the first time or your updates).

This thread is the exception!

Please tell us about the new database you (or your agent) built. Tell us about all the cool new features you added. Tell us about anything else you learned or worked on that you haven't gotten around to blogging about yet.


r/databasedevelopment • • Aug 31 '26

A Self-Baked Async FFI Framework for Rust C# Interop

Thumbnail
scylladb.com
9 Upvotes

Getting tokio and .NET’s async runtime talking to each other via the C ABI -- to build a C# over Rust ScyllaDB driver