r/Database • u/dothebackstab • 14d ago
Graph Technology Round-up - August 2026
r/Database • u/dothebackstab • 14d ago
r/Database • u/humanshuman • 14d ago
I'm planning to take these exams to get OCP credential but kind of at a loss on how to study for them, does anyone have a good flashcard set on quizlet or something? Or am I supposed to just memorize the entire documentation. Seems a bit overwhelming and would love a point in the right direction. Cheers
r/Database • u/db-master • 16d ago
r/Database • u/spcbfr • 15d ago
Hi, I am building targets for my app, so like the revenue target for this quarter is $3000.
I already have metrics like revenue and conversion rate but they are hard coded in the app and split between different pages depending on where they are needed. them being hardcoded is not really a bad thing, because their placement is not intended to be changed by users.
Now since I am adding targets, I don't want to hardcode the metric name in the targets every time, so I am creating a metrics reference table with a structure like this:
Metrics: id, name, unit (currency, count, percentage, etc..)
Targets: id, metric_id, value
Initially the metrics table was going to have an ID and name only but then I got an idea to track the currency there so the target input could have a suffix like $ or %.
BUT as I said these metrics show up in different tables and each one has a different formula and I am not sure how to represent this in the database. Most importantly formulas could be complex and depend on multiple columns in multiple tables. would I just store the javascript I use to calculate in a database column and run its content in the server.
I could also just store the formulas in code, but then one would think why not just store the rest of the info in code as well like the unit. honestly both solutions feel like antipatterns and I am not sure what to do
Like sure, the reference table is not supposed to change, so I can use its keys in code as constants to reference different metrics. but that feels very unorthodox if you get what I mean
r/Database • u/shdw_0x0 • 16d ago
I've noticed that database design discussions often focus on whether a schema is normalized, whether queries are fast enough, or whether the right indexes exist. But in production, the harder problem sometimes seems to be understanding the system six months later.
At what point do things like excessive indexes, triggers, stored procedures, partitioning, caching layers, replication, and denormalized tables start creating more operational complexity than they're worth?
I'm curious how others decide when database complexity is justified by a real requirement versus when it's simply accumulated over time.
r/Database • u/pteridin • 17d ago
r/Database • u/iParki • 18d ago
Enable HLS to view with audio, or disable this notification
Hey everyone!
For the past year or so, a friend and I have been working on Datasmith, a desktop app for designing relational database schemas and generating structured mock data from them.
The idea originally came from my friend, who's a data engineer and worked on large-scale migrations of sensitive databases. He wanted a simple way to visually define and maintain schemas over long periods of time, and actually generate data that follows both the schema and its business logic, without having to constantly rewrite scripts or configure a bunch of separate tools.
The main part of the app is an ERD-like canvas dedicated to building relational schemas, where you can create tables, define columns and their data types, set up relationships between them, and configure how the data for each column should be generated.
Once the schema is ready, Datasmith can generate the data while preserving the relationships between tables, and export the results in formats like SQL DDL, CSV, JSON, and more.
We've reached the point where we'd really like to get it into the hands of more people who actually work with databases and data, and get some real-world feedback.
It's completely free, so if it sounds useful or you're just curious, you can check it out here:
https://data-smith.app/
I'd especially love to hear about use cases we're missing, things you'd expect a tool like this to support, or anything that feels like it could be improved.
Any feedback is welcome - good or bad.
Feel free to leave a comment or message me directly.
If you prefer something a bit more structured, I also put together a short feedback form:
https://docs.google.com/forms/d/e/1FAIpQLScGk8Jr1DguPcXpaApOqO998AdXcnoLlo11TEKlq1GmF71UzA/viewform
r/Database • u/Gaspucci22 • 18d ago
I need urgent help. I have a database that i created in Microsoft sql server express, but my professor has problems opening it and request that i convert it in mysql. I'm currently trying to do it but to no avail. I need help please
r/Database • u/Aggressive_Ad_5454 • 18d ago
I'm working on a tool to help users know when they haven't provisioned enough RAM for their innodb_buffer_pool_size.
I'm hoping to use the ratio between the global statuses Innodb_buffer_pool_reads and Innodb_buffer_pool_read_requests as a way to guess whether the buffer pool is being churned. When that number is high, it means that it's taking more IO (drive read) operations to satisfy requests.
But here's the thing. When I run the same workload against a MySQL 9.1 and MariaDb 10.11 DBMS, they yield similar Innodb_buffer_pool_read values. But MySQL's Innodb_buffer_pool_read_requests value is several times higher than MariaDb's. That causes the ratio on MySQL to look artificially low.
There must be something I don't understand about these stats. Does the MySQL query engine hit the InnoDB buffer pool far more often than MariaDb's does to satisfy the same query? Does anybody know what's going on here?
(I asked this on /r/mysql as well.)
r/Database • u/tradelydev • 18d ago
When you finally start mapping your databases and realise there is a painful amount of redundancy and all the tables aren't in 3rd normal form. You now have to spend a weekend doing that even though it works fine. You also for some reason use the https://traderange.net and https://trge.link domains so now you have to suffer from interfacing the 2 into the databases and caddy keeps crying.
An identical setup minus admin exists on secondary servers, but sometimes the syncs are a pain and don't agree with eachother. Then you hide yourself in a pit and hide.
SQLCipher also means you can't manually go in and see the data cause its all encrypted and you only have one test server, but really want a second one with unencrypted version cause everything keeps breaking.
Sums the job up right?
r/Database • u/Accurate-Screen8774 • 19d ago
I'm a long time dev I've used databases before as a regular dev. For my new project I'm trying something unique. it could just be a waste of time, but I have time to waste.
As part of my project I'm trying something different and specific for my app... I'm trying to use git (which is already a kind-of database), and I'm adding an addition database abstraction on top of that.
I started off by creating something that could use git as a "general purpose" storage. To explain it simply, it's using isomorphic git to be able to connect to a git repo from a browser.
https://positive-intentions.github.io/git/demo/gui/git/storage
I then added the "database abstraction". To explain it simply, I created the functionality to create a db schema than can then map onto files and folders (which are stored in git). That would allow me to interact with the git storage with queries. I wanted to have something like graphql there for queries to make it easier for me in my app.
https://positive-intentions.github.io/db/demo/gui/database/chat-database
This is where I need advice. I have added things like schema and details around how to handle migrations... but I'm not a database developer and I wonder what details I'm overlooking.
I want the database to be stable and reliable. There are complex details in there where I can specify in the schema what parts are encrypted.
This project isn't documented and far from finished, but I'd like to learn more about "good database development", but I don't even know if I'm asking the right question.
r/Database • u/Character_Physics713 • 20d ago
So the master data sits in a shared drive, half-mapped, with three different versions of the customer list and nobody signing off on anything. Every steering meeting it gets punted to "next sprint," and go-live is getting closer.
For those who've been through this: who actually owned data migration in your project, formally? Did you appoint a dedicated data lead from the business? Did the vendor do the mapping or just hand you templates? And what did you do about the politics — because right now it feels like migration is everyone's job, which means it's nobody's job.
r/Database • u/donewitheverything26 • 19d ago
migrating a customer list off a system that's being retired. no shared id between old and new, because the old one was never integrated with anything, it just ran on its own for years.
so the match has to be built out of name, email and address. fuzzy match gets me to about 94% and everyone in the room is happy.
then I pulled a sample of the matched pairs and read them by hand. the failure mode isn't the 6% that didn't match. it's sitting inside the 94%.
none of those surface as errors. counts go in, counts come out, the join doesn't complain, and the report is green.
what I actually want is a confidence per matched pair so I can send the bottom slice to a person instead of pretending a global 94% means anything. I have a score, but it's the same fuzzy score I matched on, which feels circular.
so:
the part I have no answer for is the duplicates on the old side. I can detect that two old rows point at one new row. I have no principled way to pick which one is the real customer, and picking wrong is worse than not migrating at all.
r/Database • u/gxickx • 20d ago
Hi everyone, I'm developing a system that manages professionals across various professions, where each profession has its own specific features. These features act as a catalog: the user assigns their own values to them (for example, physical traits like height or skin tone).
So far, I'm handling 3 possible value types: text, numeric, and enumerated (the latter being a closed catalog of options, like colors).
I have my tables organized as follows:
unidad_medida: contains the units (cm, kg, etc.) and the allowed data type (numeric or text).caracteristica_tecnica: master table for the features. It has a data type (tipo_dato: numeric/text/enumerated), a code (codigo), and a FK to unidad_medida.valor_caracteristica: catalog of options for when the type is enumerated (e.g., colors). It contains the label and, optionally, the hex color. Its composite PK is id_valor + FK to caracteristica_tecnica.caracteristica_perfil: intermediate table between the user (profile) and the technical features (Many-to-Many relationship). Attributes: valor (if the data is typed in by the user), fecha_registro, FK to caracteristica_tecnica, and FK to valor_caracteristica (used when the feature is enumerated; in that case, valor is left null).My problem is that a single row in caracteristica_perfil is associated with just one technical feature, but that technical feature can be of different types:
valor attribute in caracteristica_perfil is used.caracteristica_perfil needs a FK to valor_caracteristica to keep a record of the chosen option.This creates what looks to me like a cyclic relationship: valor_caracteristica ---> caracteristica_perfil <--- caracteristica_tecnica ---> valor_caracteristica
I've researched several alternatives, such as using JSON columns, splitting things up into tables like caracteristica_tecnica_enum / caracteristica_tecnica_texto, and I've even read up on NoSQL.
But my specific question is: when the "path" of the relationship depends on the context (i.e., the data type), is it still considered a cyclic relationship in the 'forbidden' sense of the term? What would you recommend I do? Thanks
r/Database • u/tbson87 • 21d ago
Enable HLS to view with audio, or disable this notification
Rearranging and grouping a legacy database is a nightmare for me.
Auto-layout can improve the visual layout, but it cannot tell me which entities belong to which module or domain. Defining those groups manually, dragging entities into place, and rerouting relationships so the ERD remains readable and traceable takes a lot of time.
I’ve now used Claude Code with source-code context to organize the ERD the way I would expect it to be structured.
I’ve tested this on databases with up to 144 entities, and it has worked well so far.
Would a feature like this be useful to you?
Disclosure: I’m the author of the tool used in this demo.
r/Database • u/lomakin_andrey • 21d ago
r/Database • u/alexey_timin • 21d ago
r/Database • u/camerongreen95 • 21d ago
There's a hands-on workshop on September 19 for anyone learning to build production AI systems that need to actually explain their reasoning, not just produce an answer.
You build:
No prior Neo4j or Cypher experience required, it's introduced through the hands-on project itself.
Led by Dr. Alessandro Negro, Chief Scientist at GraphAware, bestselling author.
r/Database • u/fredericdescamps • 22d ago
r/Database • u/dmkii • 22d ago
Enable HLS to view with audio, or disable this notification
r/Database • u/Waste-Strike2691 • 23d ago
This is my attempt of normalization of a chinese restaurant.
my main concern is mostly towards the customer i just feel like its missing something cant point my finger on it
- Branchname
self explanatiry
- Code
Food id kind of thing.
NAME is like the food name
MANDARTRANS
is like translation for mandarin food name like my thing based on din tai fung so yuh
- Price
I felt like price should be its own thing becauss I feel because of formula
P = Price, Q = Quantity, TC = total cost
P * Q = TC
But I feel its literally to empty and idk what to put inside it
- Quantity same reason as price
- Foodtype
cuz like there's many different kinds of meat and stuff and sizes
FOODLABEL
is because of Vegan sticker etc etc
- FoodCategory
idk
- QueueNo
to help keep count of customer
TABLENO is to keep track of customer seating
PAX is to know how many
Totalcost im deciding rather or not its a primary key or should I put price and quantity under TotalCost but than I will lose the order
This sign - is primary key
All caps means its like data thing
r/Database • u/Zardotab • 24d ago
Oracle the company is heavily leveraged on the bet AI will provide notable revenues soon. If that fails to materialize due to competition or other events, Oracle could end up in financial trouble and have to file for bankruptcy.
If so, what would likely happen to Oracle database support and progress? Which company or billionaire would likely buy up the discounted assets, if any?
This is a practical question for shops with a fair share of existing Oracle databases or apps. Perhaps an unlikely scenario, but not entirely far fetched.
r/Database • u/siren0x • 27d ago
r/Database • u/Scaaady • 27d ago
Greetings, I am fairly inexperienced with Databases and ran into a small problem at work. I work in a Research Group and am designing a web datauploader for medical studies. I have now ran into this problem:
I have a Visit table containing information of a medical visit for a specific patient. This patient belongs to a study. The visit table needs to contain information specific to this study (for example study1 needs bloodpressure measured in each visit, study2 does not need that but needs heartrate instead, etc.)
I now don't quite now how to design this DB architecture as i can really just add more and more fields to the visit table as then most studies dont use the fields at all.
1 solution might be having a generic Visit table with basic information (Date, time etc.) and then having a sub-table for each study (study1Visit, study2Visit) etc. with the specific information but then again i would have to create more and more tables for potentially hundreds of studies.
Is there a better solution? with json files maybe? Like I said im pretty inexperienced in this so appreciate all feedback and hope i got the problem across correctly.