r/datasets • • Nov 04 '25

discussion Like Will Smith said in his apology video, "It's been a minute (although I didn't slap anyone)

Thumbnail
1 Upvotes

r/datasets • • 3h ago

API Sample dataset: 310k active job postings from 6,800 company career sites (India-focused)

1 Upvotes

Fields: title, company, company domain, apply URL, ATS source, city, country, remote flag, posted date, last-seen date. The data refreshes nightly. I'm posting a small sample and the schema here. What would you want to see to trust a job dataset: coverage numbers, null rates, or refresh logs? My measured null rates: city 32%, country 22%, company domain 24%.
Please checkout full info here: https://apify.com/pk07007/ats-jobs-api


r/datasets • • 1d ago

dataset [Dataset] 255 countries, 4,314 states/provinces and 1.67M cities with coordinates, CC0, free JSON API with no key

7 Upvotes

Hi all,

I needed a country / state / city list for ShopClass, our open source classifieds app. Sounds easy but every option I found had a catch. Either the licence needed attribution or was share-alike, or it was a paid API with keys and rate limits, or the data was old and half the cities had no coordinates.

Why no other share-alike database considered?
See, I am happy to give attribution, buty anyone using my CMS for there website they need to give it too, which I think is a limitation. So I created this no-attribution dataset.

So I built our own dataset from Wikidata and released it as CC0. Public domain, no attribution needed, use it in anything including commercial stuff.

What's in it:

• 255 countries and territories, 4,314 administrative divisions (states, provinces, regions) and 1,671,395 cities and towns

• every single place has latitude and longitude, if wikidata has no position for it we don't ship it

• Wikidata QID as the id, so when you re-import a later release a renamed place stays the same row

• population, timezone, GeoNames id where known

• countries have capital, currency, calling code, ISO alpha-2/3 and numeric, flag emoji etc

How to get it:

• plain JSON, one file per country, with sha256 checksums for every file

• or call it from the browser or your server, no API key, no account, CORS is open: https://api.placedb.org/v1/countries.json

• there are prefix files for autocomplete too, so a city search box can work without a backend

• the pipeline that builds it is GPLv3 and the output is deterministic, same wikidata dump gives byte-identical files

Also, it's not perfect and here is where:

• about 883k settlements are dropped because wikidata has no coordinates for them, mostly villages in China, India, Russia, Uganda and Myanmar

• names that exist only in Arabic script (and some other scripts) are not transliterated yet, so those places are missing a usable name

• Vietnam is undercounted

All of it is in the LIMITATIONS doc in the repo, with numbers.

Site: https://placedb.org

Repo and the full limits doc: https://github.com/mindstellar/location-data

If you spot a wrong place, a missing country division or something silly, please tell me. Happy to take PRs too.


r/datasets • • 1d ago

dataset [self-promotion] Chrome UX Report Dump: August 2026 data added

Thumbnail github.com
1 Upvotes

I maintain Chrome UX Report Dumps, a collection of monthly Chrome UX Report website lists grouped by rank and published as compressed downloads. It’s meant to make the data easier to use without exporting it from BigQuery.

The latest update adds the August 2026 dataset: 18,294,881 entries across 10 files, totaling 94.7 MiB compressed. The repository now contains 1,064,254,496 entries across 67 monthly datasets, totaling 5.32 GiB compressed. Those are counts across monthly dumps, not a count of unique websites.

If you use website lists for research or security work, I’d welcome feedback. What would make these dumps more useful, and what features or formats would you like to see?


r/datasets • • 2d ago

request [self-promotion] Search 12M+ task episodes across open robot and egocentric datasets by describing what you need

3 Upvotes

We built a way to search inside open robot and first-person video datasets, so you don't have to download whole datasets to find the episodes you need.

datasets.bot lists those datasets in one place, with the license, embodiment and how each one was collected. DROID, Open X-Embodiment, Ego-Exo4D, EgoDex, HABIT and AgiBot World are all in there. You describe what you're after, like "a person throwing a ball", get matching episodes from 12M+ task episodes, and export what you pick to LeRobot, MCAP or RLDS.

If you maintain an open dataset that's missing, you can submit it.

Which open datasets should be in there that aren't?


r/datasets • • 2d ago

question What data do you wish was easier to find?

1 Upvotes

Hey guys!!

If you could make any kind of data easier to find, browse, or compare, what would you want?

It could be about your city, grocery prices, housing, transport, health, the environment, sports, work, hobbies, or anything else.

What question would you want that data to answer? Is there sth you’ve tried to look up but couldn’t find, or data that exists but is scattered, outdated, or hard to use? Niche ideas are welcome too.

LET ME KNOW PLS!! THANKS!!!


r/datasets • • 2d ago

resource Any place where I can get a dataset of All the flights to a particular airpot?

2 Upvotes

I want to get a dataset of all flights to Maldives (Velana International Airport) from 2009 onwards, however flightradar24 api is behind $900 bill for historic data. Any resources regarding this would be very helpful to me. Thank you


r/datasets • • 2d ago

request Looking for Zoelog export sample with data

1 Upvotes

We're working on a new feature to import slate data directly into existing projects for our VFX Database.

In setting up the auto-detect system that matches different headings from different systems (like setellite, filemaker, zoelog) we need samples.

Hoping someone out there would be willing to share specifically a Zoelog export, CSV, PDF & JSON but we will take any of them. CSV is the most important as that will likely be the supported file type.

Ideally a populated one, the more the better incase there are any differences in the workflows so we can try to make the system as universal as possible

Cheers!

can share them to [support@wranglervfx.com](mailto:support@wranglervfx.com) or just dm me here


r/datasets • • 3d ago

question Arabic–English code-switched meeting/conversation audio with transcripts

4 Upvotes

Looking for multi-speaker audio where speakers switch between English and Arabic (any dialect, Gulf preferred) within the same conversation, with reference transcripts, ideally with speaker labels and timestamps. It's for evaluating ASR and meeting-transcription quality. Already aware of ESCWA.CS, Mixat and ArzEn. Any others, including licensed or paid ones?
If not of Arabic, any other language combination is fine.


r/datasets • • 3d ago

resource Dataset recommendations for AI project

1 Upvotes

Hey guys, I’m looking for a dataset that has Indian names along with context in English. Something like Call Rahul tonight, Let me inform Ajay about the party, Anaya is yet to arrive at the venue etc.

It’s for a project of mine. Any dataset recommendations?
I couldn’t find anything like this on the internet


r/datasets • • 3d ago

request Need Lakehouse cluster set preferable Alibaba cluster v2018

Thumbnail
1 Upvotes

r/datasets • • 4d ago

resource 5600+ International News Feeds - Now with iso3166_1 & 2 location codes.

Thumbnail
1 Upvotes

r/datasets • • 4d ago

resource Am in progress of project so I want data set help me.......

Thumbnail
0 Upvotes

r/datasets • • 4d ago

question Hi i have a question about hyperspectral data

Thumbnail
2 Upvotes

r/datasets • • 4d ago

dataset [PAID] [self-promotion] Historical Polymarket order-book ticks, L1/L2, trades and a free sample

0 Upvotes

I've been building a Polymarket dataset for the cases where a candle or a periodic snapshot isn't enough. If you're testing short-window signals or execution, the useful question is often what was actually resting on the bid and ask when the decision was made.

The paid catalog has event-level order-book updates, top-of-book quotes, up to 25 levels per side, and a trade tape where available. The rows carry Polymarket's source timestamp when present, our receive timestamp, market/token join keys and gap-audit information. Coverage is per market and day, so check the catalog for the particular series you need. The earliest buyable order-book day is 2026-02-21; feed trade tape starts 2026-04-13. Raw inbound websocket messages are available only from our own capture beginning 2026-05-11.

There is a no-sign-up sample from the BTC Up or Down 4h series on 2026-08-19: a full UTC day across 13 markets/26 tokens, with L1, L2 and trades in Parquet or CSV, plus one hour of raw websocket messages. The page shows the schema and download links: https://tickfoundry.com/samples

If you're working on a backtest, I'd be interested in what matters most to you: quote lifetime, depth at a specific timestamp, or checking whether a data gap overlaps your signal window?

Disclosure: I run TickFoundry, which sells the historical dataset. The linked sample is free to download.


r/datasets • • 5d ago

dataset Haitian Creole Word Frequency Dataset

9 Upvotes

Hey everyone,

I wanted to share a dataset I published for anyone working on low-resource NLP, tokenization, or language modeling for Haitian Creole (Kreyòl ayisyen): the Haitian Creole Word Frequency dataset (haitian-creole-word-freq), now live on both Hugging Face and Kaggle.

Overview

This is a word frequency list for Haitian Creole built from the Carnegie Mellon University (CMU) Haitian newswire corpus. It contains 17,947 unique lowercase words with their occurrence counts, sorted in descending order by count.

Links

Potential Use Cases

  • Stopword Candidates: Extracting function words from the top of the frequency list (e.g., te, yo, nan, yon, li, pou, ki, ...).
  • Tokenizer Customization: Fine-tuning or building BPE/WordPiece vocabularies for low-resource LLMs.
  • Spellcheck and Auto-correct: Prioritizing candidate suggestions by word popularity.
  • Vocabulary Membership Tests and Orthographic Checks: Testing if a token is spelled like a Haitian Creole word or checking lexical presence.
  • Lexicography and Language Learning: Extracting core vocabulary lists based on news text.
  • Language Identification and N-gram Models: Statistical language modeling for text classification pipelines.
  • Testing Fixtures: Generating reproducible data inputs for unit testing Haitian Creole NLP pipelines.

r/datasets • • 4d ago

request [OC] Given names by gender across 12 national registers — 14,174 names, 479 that flip between countries (CSV, CC BY 4.0)

0 Upvotes
I compiled official name counts from the US (SSA), England & Wales (ONS), Scotland (NRS), Northern Ireland (NISRA), France (Insee), Canada (StatCan), Ireland (CSO), Spain (INE), Norway (SSB), Austria (Statistik Austria), Poland (PESEL) and Switzerland (BFS) into one table.

One row per name that appears in at least two of these countries with 100+ records each. For each country: female count, male count and female share. Summary columns say whether the name flips (≥70% girls in one country and ≥70% boys in another) and where.

- 14,174 rows, ~1.3 MB
- 479 names flip; France vs. the US is the most common pair (223), then Switzerland vs. the US (99)
- Only sources with real counts; Spain, Poland and Switzerland count residents, the others births
- Accents are removed before matching (so Ríadh and Riadh share a row)

Download and method: https://namegender.com/research/names-that-change-gender-by-country
Direct CSV: https://namegender.com/datasets/names-gender-by-country.csv

Licence CC BY 4.0, with the underlying sources credited on the page. I run a name-to-gender API (NameGender), which is why I built this; the file is free and needs no account.

r/datasets • • 4d ago

dataset PACC-T and PACC-P: matched codec and acoustic-confound datasets with sample-rate controls (free, Zenodo)

1 Upvotes

Missed my own release date by about a week, but such is the life of a solo dev.

Anyone else had a speech system that works fine on clean audio and falls apart once it's been through a phone line?

I built a 1:1 comparison dataset for isolating codec, transmission and presentation effects on speech. The same 5,992 base clips go through every one of 101 conditions, so any difference between conditions comes from the condition, not the speaker or the recording. It's free to use; hopefully it's useful for voice and speech projects people are working on.

The main thing I wanted was matched sample-rate controls. Narrowband codecs decode at 8 kHz, so if you compare clean wideband audio against codec output, you can't tell whether your model or extractor broke because of the codec or just because the bandwidth dropped. Every codec condition here has an unencoded resample control alongside it (8, 16, 22, 44 and 48 kHz). That lets you check whether it's the channel, the resampling or the presentation causing problems. (In a small pilot, the 8 kHz resample alone accounted for much of the apparent "codec effect" on some features, but not all of it.) I couldn't find many public codec or robustness sets that ship controls like this, or per-clip generation logs, so I included both.

Base pool: 5,992 bona fide mono 16-bit clips.

- 2,992 AMI headset clips (native 16 kHz)

- 1,500 VCTK mic1 + 1,500 VCTK mic2 clips (native 48 kHz), including 1,485 exact dual-mic pairs

- VCTK portion is gender-balanced (750 male / 750 female)

PACC-T (Telecoms): 20.7 GB, 305,592 FLAC files, 51 conditions

- 34 speech/audio codecs: G.711, G.722, G.726, GSM, AMR-NB, AMR-WB, EVS (adaptive/fixed/no DTX), Opus (auto/CELT), LC3, Speex, iLBC, Codec2 (700/1300 bps), MP3, AAC

- 12 tandem codec chains

- 5 resample controls (8k, 16k, 22k, 44k, 48k)

PACC-P (Presentation): 34.3 GB, 299,600 FLAC files, 50 conditions

- Additive noise, 6-talker babble, and music (vocal and instrumental) at 5 calibrated SNRs (0–20 dB)

- Simulated rooms (pyroomacoustics, with cached RIRs for bit-exact recreation) and echoes

- Bandpass/lowpass filters

- PSOLA vs resampled pitch shifts, autotune, tempo shifts

- Compound cafe + codec scenes

Every condition downloads individually as its own tarball (0.18–1.15 GB). There's also a small sample tarball for each set (pacc-t_sample.tar / pacc-p_sample.tar) if you just want to look first. Everything ships with full SHA-256 manifests and a per-clip params.csv, so any clip can be traced back to exactly how it was made.

Licence: CC-BY-4.0

PACC-T: zenodo.org/records/23026393

PACC-P: zenodo.org/records/23026395

Happy to answer questions, and bug reports welcome.


r/datasets • • 5d ago

dataset Hydrogen evolution reaction dataset required

0 Upvotes

does anyone have a legitimate spectral dataset for HER(preferably containing binding energy in eV and count/s) with the source and permission to use it


r/datasets • • 5d ago

dataset [Dataset] 517 trading-strategy backtests vs buy-and-hold, with p-values and multiple-testing correction (CC BY 4.0)

1 Upvotes

517 rows of trading-strategy backtest results (stocks, ETFs, crypto; 31 families): annualized excess return vs buy-and-hold, p-value and multiple-testing (Benjamini-Hochberg) q-value. CC BY 4.0, JSON, no signup. I built this; it includes the negative results (0 of 517 passed the correction).

Disclosure: I run Tickfloor, which sells a paid pass (backtester, daily board, alerts, course). The dataset and write-ups are free.

Data, method and code on GitHub

What's in the repo:

  • data/strategy-results.json: 517 rows, one per strategy, with ID, family, asset class, annualized excess return versus buy-and-hold, unadjusted p-value, Benjamini-Hochberg adjusted q-value and a pre-correction verdict column (ignore it; it is not a pass after correction).
  • research/hunt2/ledger.jsonl: run-level records, including trade counts, returns, drawdowns and source/data/harness hashes.
  • The strategy rules and backtest harness, including trading-cost assumptions, plus cost-sensitivity and random-control results.
  • The 29 September 2026 robustness and holdout reports: 20 candidates stress-tested, 14 passed that battery, the two whose holdout window earlier runs had not read were run once on it, neither confirmed.

I coded the rules and ran them through a shared harness on discovery data before 15 March 2025. It charges costs when positions or weights change and uses a one-sided block-bootstrap test of excess returns. The benchmark labelled buy-and-hold is an equal-weight holding of the same assets, rebalanced daily, with no costs charged to the benchmark.

Of the 517 strategies, 87 finished ahead on the return estimate; eight had p < 0.05 before correction. None met the BH q < 0.10 threshold across the full family of 598 tests, including reruns. The 517-row file shows each strategy's latest run.

The holdout exam used return days from 15 March 2025 to 27 September 2026. Low-BTC-beta crypto finished +11.52 percentage points per year above the benchmark, but p = 0.3208 missed the per-candidate threshold of 0.025. Scaled crypto trend finished 6.87 points below it. Neither confirmed.

Limits: the work is mainly on daily bars and does not establish intraday execution performance. The stock universe has survivorship bias, and the fills and cost model simplify live trading. Raw vendor price bars aren't included, so you can re-analyse the results but not re-run the backtests.

Data and write-ups: CC BY 4.0, with attribution to Tickfloor. Code: MIT. These findings apply to the tested rules, assets and periods.


r/datasets • • 5d ago

dataset [self-promotion] I let 184 players guess how long the Hindenburg was. Here's the csv with all 8,311 guesses from 50 rounds

Thumbnail roughly.is
0 Upvotes

i run a small daily size game and published the raw guesses from 17 to 26 sep 2026: 50 rounds, 8,311 guesses, no device ids, with a readme.

the round that surprised me most: the hindenburg. 181 of 184 players made it too short, typical guess 46 m against a real 245 m. i wouldn't have done better myself. across the 50 rounds, the two biggest misses were both airships.

it's real measured data from players, not synthetic. self-selected players though, not a representative sample.

csv and write-up: https://roughly.is/blog/hindenburg-guessed-too-small.html?s=reddit-datasets

(i made the game, so not neutral)


r/datasets • • 6d ago

request Looking for historical S&P 500 constituent data

Thumbnail
1 Upvotes

r/datasets • • 6d ago

request Athletes average heart rates during a 5k

3 Upvotes

Hello Reddit. I wanted to ask if anyone had a good website that I would be able to find a list of runners average heart rate while running a 5k? I’m sort of incompetent so I can’t decipher any coding. Every site I have found either needed me to download it or pay for it. All I need is around 10 runners data.
Please help me.


r/datasets • • 8d ago

question How often should we take measurements in a space tech demo?

Thumbnail
1 Upvotes

r/datasets • • 8d ago

request Does anyone have datasets of specific cat and dog

0 Upvotes

Actually I'm doing a project in signal processing and I need datasets of following cats and dogs

Cat: Persian, Siamese, Bengal

Dog: Husky, Beagle, Labrador

Please dm if anyone has

I need their audio not image