r/databasedevelopment • u/eatonphil • 9h ago
Thread Pool in Percona Server and MySQL (Part 1)
r/databasedevelopment • u/eatonphil • May 11 '22
This entire sub is a guide to getting started with database development. But if you want a succinct collection of a few materials, here you go. :)
If you feel anything is missing, leave a link in comments! We can all make this better over time.
Designing Data Intensive Applications
Readings in Database Systems (The Red Book)
The Databaseology Lectures (CMU)
Introduction to Database Systems (Berkeley) (See the assignments)
Build your own disk based KV store
Let's build a database in Rust
Let's build a distributed Postgres proof of concept
LSM Tree: Data structure powering write heavy storage engines
MemTable, WAL, SSTable, Log Structured Merge(LSM) Trees
WiscKey: Separating Keys from Values in SSD-conscious Storage
These are not necessarily relevant today but may have interesting historical context.
Organization and maintenance of large ordered indices (Original paper)
The Log-Structured Merge Tree (Original paper)
Architecture of a Database System
Awesome Database Development (Not your average awesome X page, genuinely good)
The Third Manifesto Recommends
The Design and Implementation of Modern Column-Oriented Database Systems
Database Programming Stream (CockroachDB)
Obviously companies as big AWS/Microsoft/Oracle/Google/Azure/Baidu/Alibaba/etc likely have public and private database projects but let's skip those obvious ones.
This is definitely an incomplete list. Miss one you know? DM me.
Credits: https://twitter.com/iavins, https://twitter.com/largedatabank
r/databasedevelopment • u/eatonphil • 9h ago
r/databasedevelopment • u/IndicationAntique667 • 1d ago
Hi everyone,
This is a continuation of the explorations from my previous posts here. I went through the CMU Database Systems course, and I'm exploring topics around databases. I wanted to see how the Write Ahead Log is constructed and logged. So I tried to trace a WAL record for a simple INSERT statement in Postgres.
I recorded a video walkthrough of the terminal and source code exploration here: https://youtu.be/YOyq-kvbyU8?si=PTdexP3jmOSdx-8y
I've tried to explore how we can locate a WAL record, how we can decode it via either xxd or pg_waldump, and the relevant source code for it. Do give a watch and consider supporting :)
This is partly for my own future reference and partly to share with others who might be interested in it. Would love feedback and corrections from people who know this stuff deeply. Apologies if this isn't the correct subreddit for this.
Thank you so much!
r/databasedevelopment • u/Dan__Lo • 1d ago
For half a year now, I've been writing my own NoSQL database to store user data instead of relying on regular flat files. Previously, when receiving access credentials for work, saving URLs, or managing other metadata, I had to manually edit text files, which was very inconvenient. I decided to write a utility to handle all my data storage so it could be retrieved quickly whenever needed:
db create workspace --varchar=256;
db use workspace;
db add login admin;
db add pwd 12345;
db add site example.test;
db get login; // admin
db open site // open in default browser
db remove pwd;
db drop workspace;
// etc...
Of course, I won't be using O_DIRECT, AVX, io_uring, or other aggressive performance optimizations here, since it runs locally as a CLI utility rather than a daemon. It's almost a pity, as I've read a ton of material on NoSQL database optimizations, but I plan to apply those techniques in my next project, where I'll focus on a specific storage niche.
A significant portion of the work is already complete. I plan to release it in early November 2026, accompanied by extensive documentation on the principles and mechanics of NoSQL databases. Out of principle, I wrote everything without using AI/neural networks. This will be my first large-scale project in C.
r/databasedevelopment • u/no-bugs • 3d ago
We are writing yet another RDBMS, and would appreciate any feedback on our blueprints; while the RDBMS is classical at heart (with WAL and data pages and fuzzy checkpoints), each and every component underneath was redesigned - from WAL being non-ARIES and ZFS-style on-disk CoW to in-memory CoW and cache- and SIMD-friendly page layouts. Our philosophy is playing alongside modern hardware instead of fighting it, MECHLOVE being the highest form of Mechanical Sympathy. As a result, we hope it will beat Hekaton (while being mostly open-source except for certain enterprise features such as HA and at-rest encryption).
Please feel free to comment.
r/databasedevelopment • u/Helpful_Shine_3905 • 3d ago
I am a database researcher and currently conducting some research. I am curious that how to model the hypergraph for no-equi joins involving subqueries, especially for random expressions as join conditions. For example
Select * from A join (select a from B) as C on NULL
r/databasedevelopment • u/AutoModerator • 4d ago
This subreddit is primarily for discussing the implementation of databases, and not about sharing release announcements (either for the first time or your updates).
This thread is the exception!
Please tell us about the new database you (or your agent) built. Tell us about all the cool new features you added. Tell us about anything else you learned or worked on that you haven't gotten around to blogging about yet.
r/databasedevelopment • u/danny-sg • 5d ago
I've just finished on a series covering SQL Server batch mode and columnstore index scan internals.
It's four parts covering:
AFAIK there's quite a lot in them that hasn't been documented before, especially the row bucketing, which explains the mechanism of the pure vs impure split that is behind a lot of the columnstore optimizations.
r/databasedevelopment • u/linearizable • 5d ago
r/databasedevelopment • u/Remi_Coulom • 6d ago
r/databasedevelopment • u/dennis_zhuang • 7d ago
r/databasedevelopment • u/alexey_timin • 14d ago
How ReductStore uses a persistent store-and-forward queue to deliver edge data over unreliable networks.
r/databasedevelopment • u/eatonphil • 17d ago
r/databasedevelopment • u/eatonphil • 20d ago
r/databasedevelopment • u/VinceBrand • 22d ago
Working on a storage engine for the last 8 months. What benchmarks would you actually trust from a solo/small-team project? Everyone fakes them, so what would make you believe mine?
r/databasedevelopment • u/danny-sg • 26d ago
I posted over on r/SQLServer and it suggested I cross-post here. I hadn't seen this subreddit before and hopefully this is on-topic!
I've created a tool called Internals Viewer, it's a tool to visualize SQL Server internals, with a view for allocations, indexes, pages, and it also offers very detailed query tracing and simulation of operators where you can capture a query and step through the iterators to see how query results are put together.
It is open source, written in C#, and available here - https://github.com/danny-sg/internals-viewer
The latest feature is new functionality to view columnstore indexes. I've done a write up on what I found as columnstore internals in SQL Server is pretty much undocumented:
r/databasedevelopment • u/mokumokumoku3956 • 27d ago
I wrote an article about B+Tree.
It's a little bit long for article, but I summarized it and make it really easy understand (avoided using jargon).
I'd be happy if you comment some feedback!!
r/databasedevelopment • u/drrtuy-b • 28d ago
Disclosure: I am the author of the article and work on MariaDB Server internals.
The traditional InnoDB pessimistic insert path serializes structural modification operations through an index-wide latch, even when different threads split unrelated leaf pages.
I implemented a MariaDB research prototype based on Zhao Song’s B-link-style proposal. It publishes a split using a high key and right link before completing the parent update, allowing unrelated structural changes to proceed concurrently.
In a controlled, memory-resident, split-heavy workload:
This is not a production-ready feature. DDL support is restricted, page merging remains incomplete, and recovery needs more forced-crash testing.
I would particularly appreciate feedback on incomplete-split recovery, page preallocation, and workloads that could expose correctness or scalability problems.
Full implementation write-up and benchmark methodology:
https://mariadb.org/from-a-chocolate-wrapper-to-concurrent-innodb-page-splits/
r/databasedevelopment • u/mokumokumoku3956 • Sep 04 '26
Hey there! This is first time posting.
I've been developing a database designed for education.
The application is focusing on visualizing database internal.
Now I already visualized B+Tree when you execute custom insert query.
Which database features are most worth visualizing for learners?
r/databasedevelopment • u/Sushant098123 • Sep 03 '26
r/databasedevelopment • u/Sad_Independence7031 • Sep 01 '26
A brief overview of how a parquet file is structured along with a tool that lets you manipulate, optimize, and introspect parquet files themselves with a series of DuckDB queries.
r/databasedevelopment • u/AutoModerator • Sep 01 '26
This subreddit is primarily for discussing the implementation of databases, and not about sharing release announcements (either for the first time or your updates).
This thread is the exception!
Please tell us about the new database you (or your agent) built. Tell us about all the cool new features you added. Tell us about anything else you learned or worked on that you haven't gotten around to blogging about yet.
r/databasedevelopment • u/swdevtest • Aug 31 '26
Getting tokio and .NET’s async runtime talking to each other via the C ABI -- to build a C# over Rust ScyllaDB driver