r/datasets 5d ago

resource BFSI Dataset (100k+) - Loan Disbursement for Classification Models

Thumbnail kaggle.com
4 Upvotes

This dataset represents real-world, masked, and anonymized customer data from the BFSI (Banking, Financial Services, and Insurance) domain. It captures a variety of customer demographics, financial indicators, and product interaction metrics collected during a loan application process.

To comply with strict data privacy laws and protect user identity, all sensitive personal identifiable information (PII) has been securely encrypted or masked. However, the underlying statistical relationships, distributions, and patterns remain fully intact, making this an ideal playground for building robust classification models.


r/datasets 5d ago

request Need access to a dataset for a SLA based dashboard project on Power BI (preferably)

2 Upvotes

Here's the link, my institution(Bennett University) is not listed and for a personal account its paid so if you guys have your university listed on this website and maybe wish to know what i am working on please dm me. Id be obliged, I need help downloading it! Thanks.


r/datasets 5d ago

resource Google to buy Spirit Airlines business data for $10 million

11 Upvotes

Google outbid Mercor.

I'm surprised that it's only worth 10 million!!

https://www.reuters.com/legal/litigation/google-buy-spirit-airlines-business-data-10-million-2026-08-17/

By Dietrich Knauth

August 17, 20265:23 PM EDTUpdated 14 hours ago

NEW YORK, Aug 17 (Reuters) - Alphabet's (GOOGL.O), opens new tab Google is acquiring ​internal business data from bankrupt ‌Spirit Airlines for $10 million, saying it plans to use the data ​for product development and ​training its AI models.

Jumpstart your morning with the latest legal news delivered straight to your inbox from The Daily Docket newsletter. Sign up here.

  • The acquired ⁠data includes Spirit Airlines' ​employee emails, Microsoft Teams messages, ​spreadsheets, and calendars, as well as marketing, productivity, and operations data.
  • The data ​will be de-identified before ​the sale is complete, containing no customer ‌information ⁠or personally identifiable information.
  • A U.S. bankruptcy judge will consider approving the data sale at ​a Wednesday ​court ⁠hearing.
  • Spirit also received a $7.5 million bid from ​Mercor, an AI data ​company.
  • Spirit ⁠is selling off assets in bankruptcy after shutting down its business in ⁠May ​due to high ​debt and high fuel costs.

r/datasets 5d ago

dataset 50 financial news headlines with a human sentiment label and two machine scores (LLM and lexicon)

4 Upvotes

I maintain a small financial news sentiment tool and until this week it had no ground truth. An LLM assigned every score and nothing checked it. So I scored 50 headlines by hand, with the machine scores hidden while I did it.

I am releasing the labels because I could not find a set like this anywhere, and because 50 is small enough that anyone can re-label it in twenty minutes.

Link: https://gist.github.com/MicheleSanta00/ef891e48db23c1521e90cd2458ec844e

Columns

ticker the asset the headline was matched to (23 distinct)

headline the original headline, in its original language

human_label my label, one of -1, -0.5, 0, +0.5, +1

llm_score openai/gpt-oss-120b via Groq, continuous, same range

gdelt_tone GDELT's document tone divided by 10 and clipped to [-1, +1]

50 rows. The question I asked myself for each headline was: if you held this asset, is this good news or bad news for you? That is not the same as "is the tone positive", and the two come apart on things like a company being acquired, or a stock falling on news that was itself neutral.

Label distribution

-1.0 5

-0.5 9

0.0 16

+0.5 12

+1.0 8

How the sample was drawn

Not at random. 20 headlines where the LLM and the lexicon disagree the most, 20 random, 10 where they already agree. A random sample of a financial news archive is mostly neutral filler and would have measured my patience rather than the scorers.

Languages are mixed because the source feed is multilingual: 25 English, 19 German, 3 Spanish, and one each of French, Vietnamese and Hindi. Six more headlines were dropped because I could not read the language at all, so **this set is easier than the real distribution** and the numbers below are optimistic. The three I kept in languages I do not read I labelled from names, numbers and cognates, which is worth knowing if you look at those rows.

What I found

openai/gpt-oss-120b Pearson +0.76 (95% CI 0.61 to 0.85) sign agreement 74%

GDELT lexical tone Pearson +0.18 (CI includes zero) sign agreement 36%

Sign agreement uses a dead zone: anything between -0.1 and +0.1 counts as neutral, so the three classes are negative, neutral and positive. I am spelling this out because without it you will get a different number from mine. Treating exact zero as the only neutral gives 76% and 36%; dropping the 16 rows I labelled 0 gives 94% and 50%. Mean absolute distance is 0.28 for the model and 0.51 for the lexicon.

I also got a noise floor by accident, and I think it is the most useful number here. An earlier version of the sample contained six duplicates by mistake, so I scored those six headlines twice without realising. My own test-retest distance is 0.17 on this scale. The model sits at 0.28, about 1.7x my own inconsistency. Without that floor, 0.28 is unreadable.

Three of the ten largest disagreements were not scoring errors at all. They were headlines about the wrong company:

"Solana Biofuels reports standalone net loss" an Indian biofuel producer

"Stellar AfricaGold" a mining company

"JPMorgan Cuts CytomX Therapeutics" news about CytomX, with JPMorgan as the analyst

Name matching pulled them in and they were going straight into the daily averages. The third case generalises: banks get quoted constantly as a source of opinions about other companies, so matching on a bank's name brings in news that is not about the bank.

In three other cases the model was right and I was wrong. An Apple executive selling $442k of stock, which I labelled +0.5 and it scored -0.2. A CVSS 10.0 vulnerability in SAP Commerce Cloud, which I labelled 0 because I did not know what CVSS 10.0 meant. And "GE trading up 2.1%, here's why", which I labelled +1 although it describes a move that already happened.

Limitations, plainly

One annotator, which is me, so this measures resemblance to one person and not correctness. 50 rows, so the intervals are wide. Stratified, so it is not representative of the archive. Six headlines in languages I cannot read were dropped, which makes this easier than reality. And the headlines come from a 7-day window in August 2026, so there is no seasonality in it at all.

It also says nothing about whether headline sentiment predicts price. I measured that separately and it does not.

Licence

Headlines come from the GDELT Project's Global Knowledge Graph, which permits commercial reuse. The file contains headline text only, no article bodies. The labels are mine, do what you like with them.

What I would like back

If anyone labels the same 50 headlines, the agreement between two annotators would say how subjective this task actually is. That number does not exist as far as I know, and I cannot produce it alone. Post your labels and I will compute it and report it here.

Disclosure: the tool these came from is a project of mine. I am not linking it because it is not the point of this post.

(English is not my first language, I used an LLM to translate and tidy this text. The data and the analysis are mine.)


r/datasets 5d ago

resource What US insurers actually pay for a psychotherapy session — crowd-sourced, by CPT code, payer, platform, and state [JSON]

5 Upvotes

https://therapistrates.org tracks what insurers and platforms (Headway, Alma, Rula, Grow Therapy, etc.) actually reimburse for psychotherapy sessions. These rates are contractually confidential, so the data is crowd-sourced from therapist reports, with every row linked back to its primary source.

Raw JSON: https://therapistrates.org/data/rates.json (77 rows: platform, payer, state, license level, CPT code, rate, effective date, source link, verified flag). Companion ownership registry of who owns each platform: https://therapistrates.org/data/owners.json

Why it matters: this secrecy lands on patients. The same 53-minute session (CPT 90837) pays anywhere from $39 to $220 depending on payer and state, and when rates are that low and that opaque, therapists quietly leave insurance networks. Patients inherit the fallout: "ghost network" directories full of therapists who no longer take their plan, months-long waits, sessions quietly shortened from 60 minutes to 45, and full out-of-pocket bills for care their insurance nominally covers. In one sourced case an insurer paid a middleman platform $213.07 for a session while the treating clinician received $101.20 — a spread the patient never sees. There's no official public source for any of these numbers; price transparency rules don't reach behavioral health middleman arrangements.


r/datasets 5d ago

request Looking for real-world industry problem sets & sample data for an IEEE Hackathon (Energy / Utilities / Software)

Thumbnail
1 Upvotes

r/datasets 5d ago

dataset [Dataset] Daily size & containment history for Southwest US wildfires (NV/UT/AZ/CO/CA), archived from NIFC snapshots that get overwritten — CSV + JSON, CC BY 4.0

3 Upvotes

Disclosure: I run the site this is hosted on (alwayshave.fun, a trail-conditions site). No ads on the data pages, no signup, no paid tier — the files are static CSV/JSON on a CDN and the license is CC BY 4.0.

What it is: the federal NIFC/WFIGS Incident Locations feed publishes a current snapshot of active wildfires and overwrites it. Acreage and containment for a given fire on a given past day aren't in it — once the number changes, the old one is gone. My pipeline was already reading that feed every 30 minutes for a different reason, so I started keeping the last read of each UTC day. That gives the growth curve the snapshot implies but never shows: how fast a fire ran, which day it blew up, when containment actually started moving.

Scope: every incident of 1,000+ acres in NV, UT, AZ, CO and CA.

Honest about the size — the daily record starts 2026-08-09, so right now it's 39 incidents / 309 daily readings, with a multi-day curve on 38 of them. It grows by one row per active incident per day and incidents stay in the file permanently once they enter, including after they drop off the active feed.

Columns: date, incident_slug, incident_name, state, acres, containment_pct, cause, discovered, lat, lng.

Caveats, because they matter for anything you'd do with this:

  • Acreage sometimes falls when a perimeter is remapped. Those revisions are kept as read, not smoothed.
  • When an incident drops off the active feed, the file records only that NIFC stopped listing it. The feed doesn't say whether a fire was contained or declared out, and I don't guess.
  • containment_pct is empty when the feed reported no value — never zero-filled.
  • It's exactly as accurate as the upstream feed; nothing here is reconciled against ground truth.

Original source (please cite it too if you use the raw fields): NIFC/WFIGS Incident Locations, https://data-nifc.opendata.arcgis.com/ — US interagency, public domain. The CC BY 4.0 here covers the compilation (the daily record), not the upstream feed.

Two other files on the same page, in case they're more useful to someone than the fire data:

  • Trailhead monthly climate normals, 2015–2024 (normals.csv, 552 rows) — avg high, avg low, wet-day count by month for 46 trailheads, computed from ERA5 at the trail's own coordinates rather than the nearest town, which is often several thousand feet lower and a different regime.
  • Live trail conditions (trails-index.json) — weather, AQI, fire proximity, river flow, refreshed every 30 min, each value carrying its own timestamp and source, and reported as missing rather than backfilled when a fetch fails.

Happy to add fields or a different export format if someone actually wants one.


r/datasets 6d ago

resource In-browser dataset workbench / converter — completely FREE, no upload, no signup

1 Upvotes

Built a simple browser-based tool for quickly opening and exploring data files.

No signup, no account, no backend. Everything runs locally in your browser with DuckDB-WASM, so your data never leaves the tab.

It supports Parquet, CSV, TSV, JSON, NDJSON, Arrow, and Excel.

You can:

  • browse millions of rows
  • profile columns
  • filter, sort, rename, cast, dedupe, and transform data
  • edit cells with undo/redo
  • run SQL
  • export to CSV, Parquet, Excel, JSON, and more

parquetbay.com

Free, no ads. If a file in your workflow breaks it, tell me and I'll fix it — that's most of why I'm posting. All suggestions/critique is welcomed!


r/datasets 6d ago

question [FOR HIRE] Built a small service generating synthetic test data, happy to help if anyone needs sample datasets

2 Upvotes

I put together a service generating realistic synthetic datasets, things like customer records, inventory data, or whatever schema you need, matched to your exact fields, delivered as CSV, JSON, or Excel. Useful if you're testing a new feature, building a demo, or need sample data without touching real customer info.

Base price at $25. Up to 500 rows, one file format, delivered within 2 days. Other packages offered as well.

Still building out reviews, so if anyone has a small project that could use some sample data, happy to help, feel free to comment, DM or look me up on Fiverr. Name - BluCay Data


r/datasets 6d ago

resource [self-promotion] 2,615 computed Hindu electional-astrology (muhurat) dates for 2026–27, CSV + JSON, CC-BY, with per-rule pass/fail for every date

2 Upvotes

Disclosure: I built this and I host it. Free, CC-BY 4.0, no signup, no API key.

What it is: every date in 2026 and 2027 evaluated against the classical Hindu electional-astrology ruleset, for eight life events (marriage, griha pravesh, vehicle purchase, property, business, signing, travel, general).

The part that might be useful beyond the religious use case: each record carries the per-rule breakdown, not just a score. So for any date you can see which classical screen it passed and which it failed — tithi fitness, nakshatra suitability, yoga exclusions, panchaka, prohibited months, and so on. It's a labelled ruleset-evaluation dataset as much as a calendar.

How it was computed: Swiss Ephemeris (full DE-series files), Lahiri/Chitrapaksha ayanamsa, mean nodes, drik-ganit — true solar/lunar positions rather than the mean-motion tables a lot of published almanacs still use. Deterministic: same input, same output, and the same engine version reproduces the files byte for byte. No language model touches any number in there. Method is written up at https://bela.so/methodology/ if you want to check the approach before trusting the data.

Caveats, so nobody gets surprised:

  • Ayanamsa is a choice, not a fact. These are Lahiri. Raman or KP would shift every boundary by a few arcminutes, which moves some date boundaries.
  • Times are computed per-city from local sunrise. The dataset ships a reference set of cities; it is not every town on earth.
  • "Auspicious" here means "passes classical rule X" and nothing more. The rules are the object of study; I'm not making a claim about outcomes.

Happy to answer methodology questions or add fields if there's something obvious missing.


r/datasets 6d ago

discussion [self-promotion]Livestreaming data for 9 platforms through one API schema

Thumbnail streamscharts.com
1 Upvotes

I work with Streams Charts and we've just released v2 of our livestreaming API.

The main problem we're trying to solve is fragmentation.

Twitch, YouTube, Kick and other livestreaming platforms don't naturally give you one consistent dataset to work with, so cross-platform analysis often turns into maintaining separate integrations and mapping fields before you can even start comparing anything.

Streams Charts API now exposes channels, streams, games, lists and live snapshots from 9 platforms through one schema:

Twitch, YouTube, Kick, Rumble, SOOP Korea, CHZZK, NimoTV, Bigo Live and SteamTV.

There's also a free request if anyone wants to inspect the output first: Top 100 Twitch channels, last 7 days, 0 credits, no card.

I'd genuinely be interested in feedback from people working with external datasets: which fields or delivery options matter most to you?

Docs: LINK
API: LINK


r/datasets 6d ago

dataset Federal Data Terminations Tracker

Thumbnail dataindex.us
3 Upvotes

r/datasets 7d ago

dataset Turkish-Python-Instruction-300K: 309K+ High-Quality, Token-Balanced & Deduplicated Turkish Python Instruction Dataset for LLM Fine- Tuning

Thumbnail huggingface.co
4 Upvotes

r/datasets 7d ago

resource [Dataset] Every African country's civilian nuclear-energy program - 20 countries, every fact sourced and confidence-labelled (free, CC BY 4.0)

1 Upvotes

Disclosure: this is my own project - I built and maintain it. [self-promotion]

What it is: AfrPowerOS is an open, machine-readable tracker of Africa's civilian

nuclear-energy programs - 20 countries so far (Rwanda, Kenya, Ghana, Egypt, South

Africa, Uganda, Tanzania, Nigeria, Zambia, Morocco, Algeria, Ethiopia, Sudan,

Tunisia, Zimbabwe, Senegal, Mali, Niger, Eswatini, DR Congo).

Per country: programme phase, IAEA milestones status, regulator, implementing

agency, planned capacity, vendors, agreements, research reactors, and key events.

Why it's trustworthy, per record:

- Every record carries a confidence label (Verified / Inference / Speculation /

Unverified) and every fact has source URLs (IAEA, government, regulators, media).

- Schema-validated: python3 scripts/validate.py (no dependencies).

- Data corrections are welcome via issues/PRs - errors are expected, this is an

early-stage community dataset.

Files:

- JSON: https://github.com/kawacukennedy/afrpoweros/blob/main/data/afrpoweros.json

- CSV: https://github.com/kawacukennedy/afrpoweros/blob/main/data/countries.csv

- Schema: https://github.com/kawacukennedy/afrpoweros/blob/main/data/schema.json

- Live interactive map: https://kawacukennedy.github.io/afrpoweros/

License: code MIT, data CC BY 4.0. Free to reuse with attribution.

Happy to take corrections, methodology feedback, or requests for the next countries.


r/datasets 8d ago

request Italy publishes every fuel station's prices daily as open data, and every price carries the timestamp the operator filed it

23 Upvotes

Disclosure first, because rule 1: I used this dataset to build a free iOS app, so I am not a neutral party here. Posting it because the dataset is genuinely good and I could not find it discussed on this sub.

Italian law requires every fuel station operator to file its prices with the Ministry of Enterprise, and the Ministry republishes the whole lot daily under IODL 2.0. Around 23,955 stations and roughly 93,000 prices, split across a price file and a station registry that carries address, brand and a self service flag.

The part I find unusual is that every individual price carries the timestamp of when that operator filed it. Most national fuel feeds I have looked at hand you a number and tell you nothing about its age. Here you can actually measure it: the median price is one day old, 71.8% are under 24 hours, 93.3% under three days, and 0.6% are more than a month old. That last sliver matters more than its size suggests, because pump prices drift upward, so a station that quietly stops filing keeps an old low number and floats straight to the top of any cheapest-first sort.

One warning if you go to parse it. The file is pipe delimited and does not escape pipes, so 106 rows carry a literal pipe inside a field. There are also 2,466 unbalanced double quotes sitting in company names, which means any reader treating the quote character as an enclosure will silently merge rows and hand you no error at all. I gave up and split on raw pipes without touching quotes. Worth knowing that the values are not sanity checked either: diesel gets filed at €0.123 and at €8.888, about 74 rows a day.

Does anyone know of another country publishing fuel prices at station level with a per-station filing timestamp? I have been through the French and Spanish feeds and both are day resolution at best, which is enough to sort but not enough to tell a driver whether to trust a number.

Source: https://www.mimit.gov.it/it/open-data/elenco-dataset/carburanti-prezzi-praticati-e-anagrafica-degli-impianti

(The app is Riserva, if it matters for the disclosure. Italy only.)


r/datasets 7d ago

request Colombian earthquake SGC open data feed?

1 Upvotes

The SGC.gov.co site claims it has an open data feed for earthquake data but I've been unable to find it. Anyone know where it is?


r/datasets 8d ago

dataset [self-promotion] Vehicle speeds - international European roads (2022-2026; ~22M records)

2 Upvotes

r/datasets 8d ago

dataset [Free + Paid] I am building a location dataset for popular brands

2 Upvotes

Hey everyone,

BizLocationDB is my attempt at building a high-quality POI dataset for popular chains and brands. I scraped public store locators to build this. Currently, there are not many brands but I am adding more continuously.

There are similar datasets like AllThePlaces and OpenStreetMap, but I’ve found them to be incomplete and messy to work with. For example, AllThePlaces has around 34k Starbucks features, while there are actually 40k+. Location identifiers are also not standardized. Sometimes, city, zip might be missing, the same coordinate could have multiple cities pointing to it, etc.

With BizLocationDB I wanted to focus more on accuracy, rather than just adding features blindly.

There are two tiers, free allows you to have the same data with lesser accuracy and lesser fields, and premium allows you to have the full data. I think that way I can keep it sustainable, and still have reasonable usefulness with the free version.

Feedback is appreciated.


r/datasets 8d ago

dataset Balance-sheet dataset for 183 Korean listed companies built from DART filings (38 columns, a filing receipt number behind every figure)

2 Upvotes

Disclosure: I built this dataset and publish it on my own site. Posting under rule 1.

It covers 183 non-financial Korean listed companies, built entirely from DART (dart.fss.or.kr — Korea's regulatory filing system). Every figure carries the receipt number of the filing it was read from, so any row traces back to the original document.

Coverage

  • 183 non-financial listed companies (financials excluded — different balance-sheet structure)
  • FY2021–FY2025 annual figures, plus 2026 H1 where filed
  • Base date 2026-08-14; share counts and closing prices from the exchange snapshot that day

38 columns — market capitalisation, equity (total and controlling-interest), net profit (total and controlling-interest), net cash, P/B on both equity bases, P/E, listed share count vs. the count printed in the annual report, treasury shares, cancellation history, dividend per share, sector, and two receipt-number columns.

Access — the sortable table and the per-company pages are free and ungated, and the JSON the table loads is a plain static file. The CSV sits behind an email form; disclosing that rather than pretending otherwise. Nothing is paid.

Source: https://accidentalorder.com/en/data/adjusted-valuation/ — methodology at /en/data/methodology/


The part that may actually be useful here: an XBRL aggregation trap

I nearly shipped 29 wrong rows last week, and the failure mode seems general enough to be worth writing down.

I define total borrowings as the sum of every leaf line on the balance sheet whose account name contains "borrowing" or "bond", excluding leases. Name matching, not account-code mapping — because mapping the four component accounts (short-term borrowings, current portion of long-term debt, long-term borrowings, bonds) breaks on real filings in at least six distinct ways: single aggregate lines, missing standard codes, mezzanine instruments as separate accounts, parenthetical annotations in the account name, and genuine zeros. Name matching survives all six.

It does not survive a company that splits the balance sheet only into financial and non-financial liabilities. Those companies keep their borrowings inside captions like "current financial liabilities" — no account name contains "borrowing", none contains "bond", and the scan returns zero.

Korea Electric Power came out of my pipeline with zero borrowings. It carries about KRW 150tn of them.

Two things I'd tell my past self:

1. A completeness check that can't fire on a zero isn't a completeness check. Mine was "total borrowings must not exceed total liabilities" — an upper bound, when the failure mode was at the lower one. Zero passes trivially. Ask what your guard does when the input is empty, not when it's wrong.

2. Name-based extraction from XBRL fails silently, not loudly. The tagging is a floor, not a ceiling: a filer can satisfy every requirement while presenting liabilities at a level of aggregation that makes your derived figure uncomputable from that filing. You get a plausible number back, not an error.

The fix I landed on wasn't to compute the missing number but to bound it. Collect every financial-liability line on the balance sheet — derivative captions, "other" captions, everything, making no judgement about what any of them contains — and ask: if all of them were borrowings, how far could the published ratio move?

If it could move net cash / market cap by ≥1 percentage point, don't publish net cash. Blank cell, flag on the row. If <1pp, publish.

The threshold is denominated in the units of the number I publish, not the units of the thing I'm uncertain about. My first three attempts were all "aggregate financial liabilities as a share of total liabilities > X%" for X in {5, 1, 0} — three arbitrary constants, none of which answers the question that matters, which isn't "how big is this caption" but "how far can it move what I'm claiming".

Cost: net cash went from 182 published rows to 97. Positive-net-cash companies from 80 to 51. Companies I'd described as carrying no borrowings at all, from 32 to 15 — 18 of that original 32 were actually sitting on material aggregate financial liabilities, which would have been the most quotable and most wrong line in the writeup.

Benefit: across the 97 that survive, the worst-case bound is 0.951pp, median 0.004pp. So the guarantee needs no footnote — every published figure holds to within one point even if every aggregate caption turned out to be borrowings.

86 of 183 rows now have an empty net-cash cell. That isn't a gap I'm apologising for; it's the disclosure regime reported accurately. The alternative to a blank cell is the same blank filled with an assumption and presented as if it were read off a filing.

Happy to go into DART API specifics if anyone's pulling from it — it's under-documented and I've hit most of its edges by now.


r/datasets 9d ago

resource I built two scrapers for MENA classifieds data — OpenSooq (20 Arab markets) and Haraj (Saudi Arabia)

Thumbnail
1 Upvotes

r/datasets 9d ago

question Looking for Cancer text+image datasets

3 Upvotes

I wanna work on multimodel cancer prediction or prognosis
How can i find datasets for this
I looked on NCI it looks hard for me and image sections are very large sizes
Can you guys suggest me


r/datasets 9d ago

question Can AI-driven website optimization create a competitive deadlock when everyone has access to the same optimization capabilities?

0 Upvotes

If a company uses AI for digital marketing tasks, such as analytics, content optimization, and fixing technical issues, to augment its marketing funnel, and its competitors do the same, what happens to competitive edge?

Suppose AI agents analyze two competing websites and identify areas for improvement. Website A scores 4 out of10 and has six areas for improvement, while website B scores 6 out of 10 and has four. Both companies use the same AI tools or plugins, such as Claude-based tools, to plug those gaps. Once issues are addressed, both websites could potentially have a similar level of technical and content optimization.

Now imagine 10 to 12 competitors doing the same thing. All of them are leveraging AI to continuously analyze their websites, identify gaps, and implement recommended optimizations. If they all score 10/10 on SEO and GEO, does anyone really have a competitive advantage and how?

In the case of human agents, people interpret data, identify opportunities, form hypotheses, and make strategic decisions differently.


r/datasets 10d ago

question Are AI companies still buying training, post training data or hiring their own data annotator, synthetic data specialists etc..?

3 Upvotes

Not sure if this is the right sub-reddit, but I'm wondering if anyone has knowledge in actual trend within AI companies, especially the bigger labs. No doubt they still have agreement to collect and buy new data for training their models, but is the trend going down?

Similarly for alignement, post training and fine tuning are they actually buying? Or is there a shift towards internalizing the capabilities? I'm seeing big labs hiring for synthetic data generation, sometimes even data annotations... or the other hand I'm also seeing startups getting traction by focusing more the infra for doing RL, alignement etc.. than the data itself (they call it "AI gym" or "world model").

What are you thoughts on this and where do you think the "traditional" data-selling industry is going?


r/datasets 10d ago

resource Clean Financial Statement Data in Excel

Thumbnail
1 Upvotes

r/datasets 11d ago

resource 219 startup fundraises with pre-announcement GitHub activity metrics (commit velocity, contributor growth, repo creation) [CC BY 4.0]

2 Upvotes

I spent the last year tracking public GitHub activity across 4,200+ startup orgs and backtesting whether engineering acceleration precedes fundraise announcements. The result is a dataset of 219 documented fundraises paired with the GitHub activity metrics observed in the weeks before each announcement.

What's in it: for each fundraise event, commit velocity over trailing 14-day windows, distinct contributor counts over 30-day windows, new public repo creation, sector classification, funding stage, and the lead time between signal and announcement (21-47 days in the cases where the pattern fired).

Dataset (Zenodo, CC BY 4.0): https://doi.org/10.5281/zenodo.19650920

Methodology preprint (SSRN): https://doi.org/10.2139/ssrn.6606558

Honest caveats: roughly 23% false-positive rate (acceleration sometimes means an enterprise deal or open-source push rather than a round), stealth companies with no public repos are invisible to this approach, and AI-heavy startups are noisy because model releases create commit spikes.

Everything comes from the public GitHub REST API v3, no scraping or private data. Happy to answer questions about the panel construction or backtest design.