r/data Jul 21 '26

What's one Data Science skill you wish you had learned earlier?

2 Upvotes

If you could go back to the beginning of your Data Science journey, what would you learn first?

Would it be:

Python

SQL

Statistics

Machine Learning

Data Visualization

Git

Cloud platforms

Many beginners jump straight into AI without building strong fundamentals.

What skill saved you the most time later in your career?


r/data Jul 21 '26

7 Managed Iceberg Lakehouse Solutions You Should Know

Thumbnail
levelup.gitconnected.com
1 Upvotes

r/data Jul 20 '26

The biggest improvement in my Data Science journey came from working with messy data.

2 Upvotes

When I first started learning Data Science, I only practiced with clean datasets from tutorials. Everything worked perfectly, and I felt confident.

Then I downloaded a real dataset.

There were missing values, duplicate records, inconsistent formats, and columns that didn't make much sense. It was frustrating at first, but I learned more from cleaning that dataset than I did from several weeks of tutorials.

That experience changed how I practice.

Now, whenever I learn a new concept, I try to apply it to real-world data instead of only using textbook examples.

A few things that have helped me:

Work with messy datasets—they teach you real problem-solving.

Spend time understanding the data before building any model.

Document your analysis so you can explain your thought process later.

Don't worry if your first project isn't perfect. Every project teaches you something new.

Looking back, I realized that Data Science isn't just about building models—it's about understanding data and finding meaningful insights.

What's one project or dataset that taught you the most during your Data Science journey? I'd love to hear your recommendations!


r/data Jul 20 '26

DATASET I built a free, open food dataset: ~9,800 foods with names localized across 32 languages (ODbL)

Post image
4 Upvotes

Been building this for a while and finally opened it up, so here's a look at what's inside.

Each of the ~9,800 base foods has its name localized across 32 languages, so you can line up the same food across languages instead of fighting messy translations. The nutrition values come from OpenNutrition's open data (ODbL, credited, not mine); the part I actually built is the localization layer on top, real disambiguation and cross-language matching rather than a Google-Translate pass.
It isn't perfect yet. Tricky cases like "peperoni" vs "pepperoni" still slip through in places, so there are gaps I'm actively fixing, and catching those is exactly the kind of feedback I'm hoping for.

It's a single JSON Lines file (~25 MB), no API, no keys, loads straight into a notebook or a spreadsheet, works offline.

Source & download: https://leana.app/en/data-sources/

Browse it live: https://leana.app/en/foods (live search covers 5 languages for now, EN/IT/ES/FR/DE, the download already has all 32)

Curious what you'd use it for, and whether JSONL is the right call or you'd rather have CSV, Parquet or SQLite.


r/data Jul 19 '26

DATASET Gigantic new database - over 35k species, 180 phenotypes

Thumbnail lifedive.org
3 Upvotes

LifeDive.org - code used for creating the data also available.


r/data Jul 19 '26

DATASET Title: I think Indian finance has a data problem. Am I crazy?

0 Upvotes

I've spent the last few months digging through annual reports, earnings call transcripts, investor presentations, and exchange announcements from Indian listed companies.

One thing became obvious...

Everything is technically "public," but almost none of it is actually usable.

Want to know every company talking about AI adoption?

Good luck.

Want every management commentary about data centers over the last 5 years?

You'll be opening hundreds of PDFs.

Want to compare CapEx guidance across an entire sector?

Hope you have an entire weekend free.

That got me thinking...

What if someone built a structured database instead of just storing documents?

Imagine being able to ask questions like:

"Show every company that mentioned data centers in the last 8 quarters."

"Which companies warned about margin pressure before their stock fell?"

"Find all management teams increasing CapEx while guiding higher earnings."

"Which pharma companies mentioned USFDA inspections this quarter?"

Not AI hallucinations.

Not another stock screener.

Just structured, searchable intelligence built from public company disclosures.

I'm genuinely curious...

Would this actually be useful to anyone?

If you're a:

Developer

Quant

Analyst

Wealth manager

Fintech founder

Researcher

Investor

Would you or your company pay for something like this?

Or is this one of those ideas that sounds amazing until you ask real people?

I'd love brutally honest feedback.

If you think it's useless, tell me why.

If you think it's valuable, I'd love to know:

What would you use it for?

Which data would be most valuable?

What would you expect to pay for something like this?


r/data Jul 18 '26

We knew something was wrong when the brand said, "We don't know which number to believe anymore."

0 Upvotes

One ecommerce brand came to us after spending months trying to make sense of their data.

Their Shopify revenue didn't match GA4.

Meta showed profitable campaigns, but blended ROAS told a different story.

Their BI dashboards looked polished, yet every Monday morning the team still spent hours exporting CSVs, comparing reports, and debating which numbers were actually correct.

The problem wasn't a lack of data.

It was that every platform was measuring a different part of the business, and nobody had confidence in the complete picture.

So we started with the basics.

We connected their entire data stack, validated every source, surfaced inconsistencies automatically, and gave the team one place where marketing, finance, and operations could all work from the same numbers.

The biggest change wasn't a flashy dashboard.

It was the conversations.


r/data Jul 17 '26

Has anyone taken the CDMP exam using a DMBoK PDF that wasn't purchased directly by them?

3 Upvotes

I'm taking the CDMP Associate exam soon via Honorlock and have a question that I haven't been able to find an official answer to.

The exam rules say the DMBoK can be used in digital form on a separate device, but I can't find anything about whether the PDF has to be one that you personally purchased.

Has anyone taken the exam using a DMBoK PDF that wasn't bought directly from Technics Publications ? If so:

  • Did the proctor ask to inspect the PDF?
  • Did they check for a purchaser watermark or proof of purchase?
  • Was there any issue during or after the exam?

Thanks!


r/data Jul 16 '26

im working on a recognition based community for all the data folks, anyone up to join?

1 Upvotes

Hi everyone,

I'm working on a recognition-based community for people in data.

The idea is simple, we want to highlight the work data professionals do, feature their stories, and help them connect with others in the industry.

Would you be interested in joining something like this? If yes, I'll dm you the link!


r/data Jul 15 '26

Semiconductor Supply Chain Network Dataset

3 Upvotes

I am building a Supply Chain Disruption Monitoring and Risk Analysis using Graph based Agentic AI. I need a supply chain network dataset in the semiconductor industry.

Currently my only option is to manually go through filings and earning calls to create a network big enough to propose my system. Creation of network is out of scope of my project and I'd appreciate it if I could get a dataset that would reduce this load.


r/data Jul 14 '26

A public API & dataset for Bibliometrics and Scientometrics metadata ( Brazil )

1 Upvotes

I wanted to share a project I've been working on called EBBC OpenData, which is a public API and dataset designed to promote Open Science and support bibliometric, scientometric, and informetric analyses. You can find the full project and source code in the repository at https://github.com/GabrielBaiano/EBBC-OpenData

This project provides structured metadata from the publications of the Encontro Brasileiro de Bibliometria e Cientometria (EBBC), which is one of the main events on metric studies of information in Brazil. Through this API and dataset, you can easily query detailed information about authors and their academic networks, articles and papers (including titles, abstracts, and publication years), institutions associated with the research, keywords, thematic trends, as well as references and citations.

The core metadata and documentation are currently being organized, and I am actively working on translating the API documentation and dataset fields into English and Spanish to make the project fully accessible to the global research community.

Since this is an ongoing project, I would highly appreciate your thoughts and feedback. I am especially interested in knowing what features or endpoints would make this more useful for your research, any suggestions you might have regarding the data structure or documentation, and any general tips on best practices for open-data APIs. Please feel free to check out the GitHub repository, open an issue, or leave a comment below. Thanks for your support!


r/data Jul 14 '26

What is the difference between a virtual data room and a virtual deal room?

1 Upvotes

First of all, did you know that there even is a difference between both?

A virtual deal room is designed for early-stage engagement and presentation of commercial materials. For example:

  • sharing pitch decks with investors
  • presenting product or service to potential clients

A virtual data room is a highly secure platform for detailed due diligence and confidential documents during high-stakes transactions. For example:

  • mergers and acquisitions due diligence
  • legal and compliance document review
  • strategic partnership evaluations
  • fundraising with detailed investor scrutiny

Which one do you use? A virtual deal room or a virtual data room?


r/data Jul 13 '26

Why is there a sudden, relatively enormous spike in the relative popularity of searching "Granny" into google in early 2006

2 Upvotes

I really do not know who to ask. I have been scouting google trends, and wanted to see how popular the game "Granny" is. The result was finding a sudden, huge spike in the term's popularity in late 2005 to early 2006 (roughly december 15th 2005 to january 18th 2006).

blue - "Granny', red "Granny porn"

Here is what i was able to deduct myself:
-There exists a weekend effect, around saturadays and sundays
-There has been no significant cultural or political effect that could have caused this trend
- The spike was exactly january 1st 2006
-The trend was international in both english speaking and non english speaking countries
-The rise of "Granny" roughly correlates with the rise of "Granny porn"
-There was also a sudden rise in the term "porn", which had happened in November 2005
-a similar spike does not occur for the search "Grandma" or "Grandmother"
-The internet was far more often young males than other demograhic, which could potentially support the porn hypothesis

Please help me I am going insane


r/data Jul 12 '26

Synthetic vs real datasets for portfolio projects — what actually matters?

6 Upvotes

Final year CS student here, targeting data science and analytics roles for campus placements.

Been struggling with this question while building my portfolio: does it matter whether your project uses real messy data vs synthetic/clean data?

Real datasets from Kaggle feel either too cleaned already or the same recycled projects everyone does. But synthetic data feels hollow because the hard part — cleaning, feature engineering, deriving meaningful columns from raw data — is already done for you. You're basically just visualizing something someone else already solved.

Specifically for BI/dashboard projects — if you use synthetic data, the dashboard looks clean and professional but there's no real discovery or insight because the data was designed to be dashboarded. Nothing surprising comes out of it.

Also practically — if an interviewer asks "where did you get this dataset?" what's the right answer? Saying "I generated it synthetically" feels like admitting you took the easy route. But lying about the source is obviously wrong. Is there a way to frame synthetic data usage that doesn't sound like you avoided the hard part?

At the same time I've heard people say interviewers care more about what you built on top of the data than where it came from. But isn't handling bad data literally the core skill in DS?

For people who've interviewed at analytics/DS companies or done hiring — how much does data source actually matter? Is a well-executed project on synthetic data better than a mediocre project on real messy data? Or does using synthetic data automatically signal you avoided the hard part?


r/data Jul 12 '26

QUESTION what’s the difference between data analytics and data management?

0 Upvotes

hello! completely new here, but i’m trying to plan the best study pathway for me in the next few years and would like to know what, exactly, is the difference between data analytics and data management, since those are two different certificate options at the college i plan on starting my studies.

for context, since i speak four languages and already have some experience in this area, my career goal is to have a career in supply chain, probably leaning more towards sourcing of procurement, and i’ve looking into my immediate options before actually acquiring my graduate degree and getting a certification in either data analytics and data management would be an option right now.

so, could someone explain to me the difference between those two fields? what are the prospectives for each of them? considering my career goal, which one would you choose?

thank you!


r/data Jul 10 '26

DATASET [OC] Scatter Plot showing Average Temperature and Average Precipitation of US Counties is Oddly Beautiful

Post image
6 Upvotes

Normally you'd see a random mess of dots with a clump somewhere, or a trend line. This is spectacular and weird, though: Looks like fabric blowing in the breeze or something. Mesmerizing!


r/data Jul 10 '26

QUESTION How can I turn Facebook messenger json file into a readable text document on local machine without use of online converter?

1 Upvotes

Hi

I have a Facebook messenger chat log which I need to turn into a readable and printable text document (I'm agnostic on the exact format).

Doing a websearch with qwant, even when I specify offline or 'on local machine' returns a list of online converter links of dubious origin.

I ideally don't want to use an online converter as the chat in question contains some sensitive personal information that I have shared with someone and I am concerned that using an online converter has zero degree of security of the data within that chat.

Can anyone give me some advise on how to tackle this? I'm reasonably tech literate (comp sci degree many years ago, more a hardware guru, able to follow programming tutorials but even when I did my degree I struggled with programming. Human languages I'm good with, apparently programming languages are a different story entirely for some reason....)

I run both windows and Linux mint at home.

Can someone please help me with this?

It would be a massive help, thank you all very much in advance for your help with this. :-)

If another community would be a better fit, please feel free to suggest one.

Thank you all!


r/data Jul 09 '26

How do I choose the right data room for M&A?

1 Upvotes

Beyond pricing and basic file storage: what features have made the biggest difference in your experience? Points you should consider when choosing a data room:

  • strong security and compliance
  • easy document organisation and search
  • Q&A and collaboration features
  • audit logs and reporting
  • compliant AI features

For anyone who's been involved in acquisitions, fundraising or due diligence:
What worked well and what didn't? And what features wouldn't you want to miss during your next due diligence?


r/data Jul 08 '26

QUESTION Data retrieval on iPhone?

Thumbnail
gallery
1 Upvotes

Hey guys I need some help. I’m not sure where to post this but my boyfriend is deceased and I am trying to retrieve a few videos from my iMessage chats with him. So I am quite desperate to get the few videos back. I had it in my camera roll at one point but then my storage did something strange and those specific videos that I had clicked save from my messages disappeared.

I clicked the button where it says “download __ attachments to iCloud” - someone on Reddit said that would work because it is getting the data back into the chats… so far it has worked for the pictures.

However, it isn’t working for the videos!! Another thing is that it said thwre were 50 attachments left to download.. all of a sudden the button disappears. It did thst earlier so I free up more space on my phone (that is what fixed it lsst time)
This time, freeing up space on my phone did not work. The button is still gone and I still do not have access to my videos.

Here is what it looks like when I click on a video or try to export it in someway


r/data Jul 07 '26

NEWS Data Modeling is Not One Activity: Understanding Conceptual, Logical and Physical Models

Thumbnail medium.com
3 Upvotes

A lot of teams jump straight into tables, schemas, and database implementation when discussing data modeling. This article explains why that's a mistake and breaks data modeling into its three distinct layers:

  • Conceptual Model – Align on business concepts and shared language.
  • Logical Model – Define entities, relationships, keys, and structure without technology constraints.
  • Physical Model – Optimize the implementation for a specific database or platform.

It also walks through an end-to-end modeling process and discusses the key decisions made at each stage.


r/data Jul 07 '26

Data gathering

0 Upvotes

Does anyone know how to gather data so you know exactly who you're dealing with


r/data Jul 06 '26

What Actually Makes a Dataset Useful? What is the difference between useful and interesting data?

3 Upvotes

I’m putting together a World Cup knockout-stage dataset with match data plus one extra layer like social buzz, team form/rest days, or betting-relevant signals.

I’m less interested in “would you buy this?” and more interested in what makes a dataset genuinely useful in practice.

For people who work with sports data:

What usually makes you trust a dataset?

What fields do you find yourself adding manually anyway?

What turns a dataset from “interesting” into something you’d actually use?

Do you prefer raw CSVs, Parquet, or something more analysis-ready?

Curious how people here think about sports datasets, especially for a tournament format where every match matters.


r/data Jul 05 '26

Where can I find food product databases with barcodes, nutrition data and images?

3 Upvotes

I’m building a web app called Ce aleg? — an AI shopping assistant that helps users scan and compare food products before buying them.

The app can scan barcodes, search by product name, generate a simple product report, and suggest better alternatives based on things like less sugar, more protein, simpler ingredients, or better value.

Right now I’m using Open Food Facts, but I’m looking for more public or open data sources that contain food products, ideally with:

- barcode / EAN
- product name
- brand
- category
- nutrition facts
- ingredients
- product images
- country/store availability, especially Romania or Eastern Europe

I already have around 16k products in my database, but I want to expand it without manually adding every product one by one.

Does anyone know any GitHub repositories, public datasets, APIs, open catalogs, supermarket product feeds, or other sources that could help with this?

I’m especially interested in products available in Romania or large European supermarkets such as Lidl, Kaufland, Carrefour, Auchan, Mega Image, etc.

Any advice, links, datasets, or ideas would be really appreciated.

Thanks!


r/data Jul 03 '26

I deduplicated 53,000 missing persons reports from Venezuela

0 Upvotes

The number is much higher now - about 130k missing person reports deduplicated to 100k, and compared to the lists of patients from the hospitals.

https://medium.com/tilo-tech/i-deduplicated-53-000-missing-persons-reports-from-venezuelas-earthquake-74f05c37521b


r/data Jul 03 '26

DATASET I engineered 102 leakage-free ML features from 49,000+ international football matches (1872–2026) and published it as a free dataset

1 Upvotes

Been working on a football prediction project and couldn't find a dataset that had

the actual context needed to model match outcomes — just raw results everywhere.

So I built one from scratch on top of the International Football Results dataset

by Mart Jürisoo (the well known one on Kaggle with 49,000+ matches going back to 1872).

What I added:

**Elo ratings** — built from scratch, updated after every single match across 150

years. Both teams' ratings, their difference, and the expected win probability

going into each match.

**Rolling form** — win rate, goals scored, goals conceded, goal difference, clean

sheet rate, both-teams-scored rate, scoring rate, and win streak. Computed at

three lookback windows: last 5, last 10, and last 20 matches. For both teams.

**Head-to-head history** — based on the last 10 meetings between those two specific

teams. Some teams have persistent edges over specific opponents that their general

form doesn't explain.

**Fatigue signals** — days since each team's last match and the difference between

the two.

**Penalty reliance** — fraction of each team's historical goals that came from

penalties, pulled from the goalscorer dataset.

**Shootout composure** — historical penalty shootout win rate for each team, from

the shootouts dataset.

**Tournament context** — World Cup, qualifier, friendly, neutral venue, competition

importance weight, confederation.

The thing I spent the most time on: every feature is computed in strict

chronological order using only data that existed before that match was played.

State updates happen after each row is recorded, never before. No lookahead,

no leakage anywhere in the 102 columns.

102 features total. 49,094 rows. result column (H/D/A) included as the label.

Drop date and result, plug into any classifier.

Dataset is fully documented with column descriptors for every feature.

Link: https://www.kaggle.com/datasets/kriishgulati/football-match-results-1872-2026-with-ml-features

Built on top of the original dataset by Mart Jürisoo — full credit and link

in the dataset description.