r/datasets Nov 04 '25

discussion Like Will Smith said in his apology video, "It's been a minute (although I didn't slap anyone)

Thumbnail
1 Upvotes

r/datasets 43m ago

dataset BIWI Kinect Head Pose Dataset Required.

Upvotes

Hi, I was looking for this dataset 'BIWI Kinect Head Pose'. But it is no longer available for the public, If anyone has the dataset downloaded please DM me.


r/datasets 4h ago

dataset [Self-promotion] Free ADA/USDC high-frequency market microstructure data — 20-level order book, spot + futures, 81k snapshots

1 Upvotes

Hi everyone,

I'm shutting down a crypto market research / ML project that I've been running since late 2025, and I've decided to release a free 7-day sample of the market data I collected for anyone interested in quantitative research, market microstructure, backtesting or machine learning.

The sample contains 81,579 ADA/USDC market snapshots with 97 columns.

What's included

  • 20 bid + 20 ask Level-2 order book levels
  • Bid/ask prices and quantities
  • Best bid / best ask
  • Mid price and spread
  • Latest spot trade price and quantity
  • Trade direction
  • Futures open interest
  • Funding rate
  • Mark price
  • Order-book imbalance
  • Market depth at different distances from mid price
  • Liquidity wall ratio
  • Book pressure

The collector runs on a nominal 5-second polling cycle. Because API calls and processing occur between snapshots, the actual median interval in this sample is approximately 6.17 seconds.

The original timestamps are preserved and the public sample has not been artificially interpolated or resampled.

Free download

Hugging Face:
rfab85/crypto-5s-market-data-adausdc-sample · Datasets at Hugging Face

Kaggle:
ADA/USDC High-Frequency Market Microstructure Data

I'd be genuinely interested to hear what people working with L2/order-book data think of the schema and what features you would derive from it.


r/datasets 8h ago

discussion _____finding govt datasets in India_____

1 Upvotes

has anyone ever needed govt datasets for a project and spent significant time finding/accessing it? or is it just me


r/datasets 17h ago

resource 970M+ Usenet (NNTP) text messages discussions dataset via API/WEB resource

4 Upvotes

I spent the last year gathering as much Usenet text data as possible from a large variety of resources and including from about 7000 live active servers as well to fill in gaps.

The end result is a searchable/filterable dataset acessible via web+api (free and paid options)

API : https://www.usenet-rewind.com/api-docs

Web: https://www.usenet-rewind.com

Binaries have been filtered out and range is from posts from 1981-Current and constantly populating. the site includes a histogram of distribution of message counts per year.

I welcome feedback or possible use cases. Built on Apache Solr 10, Mariadb 12 (rocksdb) and NVME disk hardware for fastest operation.

[self-promotion] [free] [paid]


r/datasets 19h ago

dataset [OC] Monthly open dataset: 1,341 live US data center job ads: advertised pay by role, plus employers and states (CC BY 4.0, DOI)

1 Upvotes

Disclosure first: the ads come from a US data center job board I run, so this is my own data and a link to my own project. Nothing is paywalled or email-gated.

What it is. Once a month I freeze every live ad on the board at one recorded UTC instant and publish the aggregates. August 2026 is the first release:

  • 1,341 live ads, 62 employers, 41 states, 254 city-state pairs
  • 769 ads (57.3%) carry employer-stated pay
  • Two CSVs: aggregates (live jobs by state, role family and employer, with shares) and pay bands (median advertised lower and upper bound by role and pay period)
  • CC BY 4.0

Zenodo, version DOI: https://doi.org/10.5281/zenodo.22261858 — the series DOI https://doi.org/10.5281/zenodo.22261857 always resolves to the newest edition.

Original source, where the method and every edition live: https://datacenterjobhub.com/data-center-hiring-index/2026/08

Method. The population is every ad live and approved at the cutoff, 2026-09-02T10:15:54Z. Pay is employer-stated only: where a job feed supplied its own estimate the ad counts as not stating pay, and ads with a floor but no ceiling are excluded from the bands. Annual and hourly are reported separately and never converted. A role x pay-period group is published only at 8 or more ranges. "Median advertised lower bound" is the median of the ranges' minimums, so it is a statistic about advertised ranges, not about earnings.

One cut that is not in the bundle yet. I ran the stated-pay rate by state against the twelve US jurisdictions that require a pay range in the advertisement itself. This was read at a later cutoff, 2026-09-03T15:10:37Z and 1,351 live ads, so it does not tie exactly to the frozen edition above:

  • 86.1% of ads state pay in the 9 mandate states that have live ads; 51.8% in the other 31
  • dropping the largest employer, which states pay in 174 of its 175 ads, widens that to 84.9% against 44.0%
  • 13 employers run 5 or more ads on each side of the line. Compared with themselves: 93.8% in mandate states, 81.0% elsewhere. So it is not only employer mix
  • of the eight largest data center markets, only California has a posting mandate. Virginia has none and states pay in 76% of 152 ads; Texas manages 41% of 340

That CSV goes into the September edition's folder so it is citable rather than just asserted here.

Limits. One board's ads, not the market: it skews to employers whose careers sites are ingested, and 62 employers is not the industry. Mandate status is as of 2026-01-30. 33 ads carry no state and sit outside the state figures. Every number above comes from a script in the repo.


r/datasets 1d ago

dataset I scraped 5.94 billion TikTok videos and 3.23 billion profiles in 3 weeks. Uploaded full dataset to Hugging Face for free. Step by step tutorial and code below.

53 Upvotes

Just uploaded the full 5.94 billion TikTok video dataset to Hugging Face. It’s fully open source:
https://huggingface.co/datasets/kuben-developer/tiktok-videos-4b

This dataset was collected using a TikTok mobile app reverse-engineering method I developed a few years ago. The method allowed me to extract billions of videos, profiles, comments and replies, hashtags, sounds, and more.

Full write-up and code here:
https://tiktok-api.seeksocial.io

Disclaimer: The TikTok app exposes 24 endpoints that can be accessed without a TikTok account, so the data itself is publicly accessible. But accessing it this way is probably still against TikTok’s ToS. Also, the full code is not free, I charge a small fee for access to it.


r/datasets 1d ago

dataset Free city-level business datasets (Berlin, NYC, LA, London, Paris, Rome, Madrid): registry-sourced company data. Feedback on the fields wanted.

3 Upvotes

Hey all, I work at CompanyData. A while back we published free business datasets for a few major cities on Datahub: https://datahub.io/@companydatadotcom (i can also share a link from our own website, just let me know in comments)

Each dataset covers companies in one city, sourced from official government registries. Current fields per company: registered address, registration number, SIC industry code, revenue, employee count, and ownership information through headquarter linkages, so you can see which companies belong to the same group. The Berlin one turned out to be the most popular by far, which we didn't expect.

No commercial angle here, downloads are free. We publish these because registry-sourced company data should be easier to get your hands on than it is.

What I'd really like feedback on is the fields themselves. Which ones actually matter for your work, and what's missing? Founding date, legal form, ownership links, NACE instead of SIC, officer counts, website domains? We decide what goes into the next batch based on what people here say, so be critical. Also open to suggestions for the next city or country.


r/datasets 1d ago

resource New massive and diverse vector datasets opened to the community

Thumbnail
1 Upvotes

r/datasets 1d ago

question What dataset is needed to train a Song Master Pro–level chord-recognition model, especially for jazz harmony?

1 Upvotes

I’m researching how to build or fine-tune an audio-to-chord-recognition engine comparable in ambition to Song Master Pro / Auralis Sound Prism.
The goal is not basic major/minor chord detection. I need reliable recognition of dense harmonic material: jazz, soul, funk, neo-soul, Brazilian music, film music, and arrangements with chords such as maj9, 6/9, m9, m11, 13, altered dominants, slash chords/inversions, secondary dominants, modal interchange, suspensions, passing harmony, etc.
Most public datasets I’ve found seem too limited: either simplified chord labels, weak annotations, or repertoire that does not really cover sophisticated harmony. In particular, I need time-aligned audio + chord labels, ideally with beat/downbeat information and a rich, consistent chord vocabulary.
My questions:
Which open datasets are genuinely useful for this level of chord-recognition work?
Are there any commercial/licensable datasets with high-quality, detailed chord annotations that can legally be used to train a model and ship it in commercial software?
Is a dataset such as iReal Pro-style chord charts, Hooktheory, Ultimate Guitar, Chordify, or similar usable in any legitimate/licensable way — or are they generally not viable due to rights and annotation quality?
For a serious model, is the realistic route to combine public datasets with a privately licensed/hand-annotated corpus? If so, roughly how many accurately annotated tracks would be needed before it becomes meaningfully good at jazz-influenced harmony?
Are there papers, benchmarks, companies, or dataset vendors I should study before spending money?
I’m specifically looking for practical, legally usable data sources—not advice to scrape chord sites. Any experience from people who have trained MIR / chord-recognition models would be very valuable.


r/datasets 1d ago

dataset [OC] Dataset of 51,080 publicly funded broadband locations with address availability results

1 Upvotes

I recently completed a dataset containing 51,080 publicly funded broadband locations in Virginia.

The project involved combining multiple public datasets, including statewide address points, administrative boundaries, broadband funding records, and service qualification results into a single address dataset.

The largest challenge was not collecting the data but reconciling it. No single source contained all funded locations. Records had to be standardized, deduplicated, geographically validated, merged across multiple funding programs, and linked back to authoritative address points while preserving provenance.

The final dataset contains:

• ADDRESS
• CITY
• COUNTY
• STATE
• ZIP
• LATITUDE
• LONGITUDE
• SOURCE_COUNTY
• FUNDING_TYPE
• SOURCE_DATASET
• VATI_INCLUDED
• BEAD_INCLUDED
• CURRENT_APB_RESULT
• APB_V2_RESULT
• APB_can_apb_service_address
• APB_lookup_status
• APB_match_quality
• APB_lookup_type
• APB_available_bundle_count
• APB_available_bundle_names
• APB_lookup_timestamp

After constructing the dataset, each address was evaluated through the public All Points Broadband ordering system and classified as:

yes
no
unknown_no_suggestion
unknown_bad_resolution
unknown_error
not_checked

Several issues emerged during processing, including duplicate addresses appearing in multiple programs, inconsistent formatting across source datasets, addresses resolving to incorrect counties, and addresses resolving to entirely different states.

Rather than discarding problematic records, unresolved locations were retained and classified separately so that every funded location remained accounted for.

The project ultimately became an exercise in address matching, GIS integration, provenance tracking, and public infrastructure data management.

Primary data sources:

Virginia Geographic Information Network (VGIN)

https://vgin.vdem.virginia.gov

Virginia statewide address points

Virginia Administrative Boundary Dataset

Virginia Telecommunication Initiative (VATI) 2022

Virginia Telecommunication Initiative (VATI) 2023

Virginia Telecommunication Initiative (VATI) 2024

Virginia Broadband Equity, Access, and Deployment (BEAD) post challenge locations

Virginia Office of Broadband

https://www.dhcd.virginia.gov/broadband

Virginia Broadband Map

https://vast.virginia.gov

Northern Shenandoah Valley Regional Commission

https://www.nsvregion.org

Rappahannock Broadband Authority Project Resources

https://rappbroadband.org/about-the-project/


r/datasets 2d ago

question Handling mixed date precision (DD.MM.YYYY, MM.YYYY, YYYY) in Excel for filtering and future Power BI reporting

Thumbnail
1 Upvotes

r/datasets 2d ago

request Open dataset for firm-level AI workforce / AI skills data?

5 Upvotes

Looking for an open-access dataset with firm-level AI workforce or AI skills data — global coverage, up to the present.

I've looked at Revelio Labs and Cognism, but all of them are paid licenses.

Is there anything open or free for academic use?

Thanks.


r/datasets 2d ago

API Title: Looking for one pilot user for a normalized SEC Forms 3/4/5 API

1 Upvotes

I’ve been building a normalized dataset and API from SEC ownership filings (Forms 3, 4 and 5, including amendments). It currently covers roughly 869,000 filings and more than 2.4 million transaction line items, including issuer and reporting-owner data, roles, transaction codes, timestamps, and price-quality information.

I’m not launching it as a public API yet. I’m looking for one person who already works with SEC ownership data and has a concrete research, backtesting, or data-engineering task they would be willing to run against it.

The pilot currently provides one filtered transaction endpoint with cursor pagination through an individual, revocable API key. There is no bulk export or separate filing/footnote endpoint yet.

The documentation is available here:

https://insiderfilings.info/api/v1/docs/

If this fits something you are already working on, please comment or DM me with a sentence or two about your use case. In return for access, I’m mainly looking for candid feedback on whether the API actually saves work, where the data model is unclear, and what blocks the workflow.


r/datasets 2d ago

request Egocentric POV / OTS Data Needed, Large Number of Hours | Worldwide

1 Upvotes

We're an AI training data marketplace working with companies building robotics and physical AI systems. We're currently looking for a data supplier, collector, or lab that has a large amount of egocentric video data and is willing to sell or license it.

Currently looking for 10,000 to 50,000 hours, with priority for OTS (off-the-shelf), already recorded material. This could be previously sold or unsold. We're interested in various tasks and regions, in both residential and commercial settings.

We have the budget to purchase, and the timeline is 3-6 months.

If you're able to produce those hours, let us know what your collection volumes per month look like. Happy to consider if your data and price are the right fit.

We're also interested in various setups (IMU, stereo, hand tracking, annotation, narration, etc.).


r/datasets 2d ago

request Anyone from institutions which provide access of dataful.in, we need data of MPLADS for our SIH project, so if there's anyone who can help...

Thumbnail
1 Upvotes

Doo heelppp guyssss plzzzz 😭😭


r/datasets 2d ago

request Hi , I’m Pia : urgently looking for a data based role

Thumbnail
0 Upvotes

r/datasets 2d ago

resource Built a qualitative analysis tool — offering free sentiment/thematic analysis to test it on real data

1 Upvotes

I built a tool for qualitative research and I need people to break it.

I’m a student with a background in ML/data science, and I’ve been building a tool that can analyze things like interview transcripts, focus groups, open-ended survey responses, reviews, etc. — basically qualitative text.

It does sentiment analysis, coding, and thematic analysis.

But here’s the problem: I don’t want to test it on fake ChatGPT-generated data.

I’m looking for researchers/students who have real qualitative data they’ve already collected and wouldn’t mind letting me run through the tool.

In return, I’ll send you the analysis/results completely free.

No subscription. No sales pitch. No “book a demo.” 😅

I’m mainly trying to figure out:

\- Are the codes actually useful?

\- Do the themes make sense?

\- Does it save you time?

\- Where does it completely screw up?

\- Would you actually use something like this in your research?

You can anonymize/redact anything sensitive before sending it.

If you have a thesis, dissertation, research project, interview transcripts, focus group data, survey responses, or anything else qualitative sitting on your laptop, I’d genuinely love to test it.

Comment “interested” or DM me.

And yes, I’m specifically looking for people who are willing to tell me “this is terrible” if it is. 😂


r/datasets 3d ago

dataset I trained a 67M-param LaTeX OCR model that runs on a laptop CPU — and built a new style-aware dataset to train it. Weights, data, and training code all open (MIT).

2 Upvotes

Hey everyone! I've been working on a little side project I wanted to share: latex-ocr, a standalone formula OCR model — you feed it an image of a math formula, it spits out the LaTeX source.

The main hook: it's only 67M parameters, so it runs comfortably on a laptop CPU. No GPU, no 300M-parameter monster to load. It's a CoCa-style model (contrastive captioner adapted for OCR), and despite the small size it beats the 107M UniMER-tiny baseline and gets pretty close to the 325M one on plain formulas.

The part I'm actually most proud of is the dataset. Real papers don't just use plain symbols — you see \mathbb{R}, \mathcal{F}, \mathfrak{g} everywhere, and existing OCR datasets basically ignore font styles, so models trained on them can't read (or hallucinate) those macros. So I rebuilt ~1.3M formulas with a MathJax → SVG → PDF → PNG pipeline and injected font-style macros with semantic heuristics (number sets → \mathbb, vectors → \mathbf, differentials → \mathrm). On that styled test set it clearly outperforms all the baselines — fair warning though, those baselines are zero-shot on styled data, so take that comparison with a grain of salt. The plain-split numbers are the like-for-like ones.

Everything is open: model weights and dataset on Hugging Face, training recipes included if you want to reproduce or fine-tune it yourself, MIT license. There's also a FastAPI server and a Gradio web UI, so you can drag-and-drop an image and see the LaTeX with a rendered preview.

Repo: https://github.com/PadishahIII/latex-ocr Model: https://huggingface.co/PadishahIIIXXX/latex-ocr Dataset: https://huggingface.co/datasets/PadishahIIIXXX/latex-ocr-dataset

Happy to answer questions about the training setup, the data pipeline, or anything else. Would love feedback — especially if you try it on your own gnarly formulas and it breaks, that's genuinely useful.


r/datasets 2d ago

API [Self-promo: I built it] Real SEC ownership-filing metadata (Form 4 / 13F) as a free JSON API — sample dataset of 23 tickers

1 Upvotes

Disclaimer per sub rules: I built and operate this — it is self-promotion, disclosed up front.

The dataset: real SEC EDGAR ownership-filing metadata — Form 4 (insider trades), 13F-HR (institutional holdings), Form 3 — for a sample of 23 major US tickers, last ~120 days (675 records currently). Original source is the SEC's public EDGAR system (public domain): https://www.sec.gov

Fields per record: ticker, CIK, form type, filing date, accession number, primary document, and the SEC source URL.

Access: free demo API key at https://edgarfeed.onrender.com (no card).

Honest limits: metadata-level records only (not parsed line items yet); sample coverage, not the full market; this is a beta demand test, not a finished product. If you need full history, webhooks, or normalized transaction line items, that's the planned paid roadmap ($0/$29/$79 — not yet verified).

Questions about the data or the approach are welcome.


r/datasets 3d ago

dataset [self-promotion] Daily dataset + free API for Pakistani mutual fund NAVs (MUFAP-sourced)

1 Upvotes

Disclaimer: this is my own open-source project.

I maintain a daily-refreshed dataset of Pakistani mutual fund NAVs, built from the public MUFAP data (the industry association at mufap.com.pk is the original source). It is served as a free no-key API.

Per fund you get the current NAV, full NAV history and computed returns, plus AMC and category metadata. JSON, MIT licensed, no PII.

https://github.com/saadsalmankhan/pakistan-mutual-funds-api

Happy to expose extra fields if they are useful for research.


r/datasets 3d ago

discussion Evidence of Fraud in an Influential Study About Procrastination

Thumbnail datacolada.org
28 Upvotes

r/datasets 3d ago

mock dataset Free High-Fidelity Dataset - Retail POS Transaction

Thumbnail github.com
1 Upvotes

A high-fidelity synthetic retail POS (Point of Sale) transaction dataset containing over 1,190,000+ synchronized master invoices and 4,190,000+ item line details across multiple relational schemas. Generated using a custom, highly-optimized Prolog simulation engine, this dataset is enterprise production-grade and perfect for database stress-testing, query optimization benchmarks, and complex retail machine learning models.


r/datasets 4d ago

request kalkine reporting data aggregator for investment

Thumbnail
1 Upvotes

Looking for information about this company


r/datasets 4d ago

request Needing a consonant zh,ch,sh,z,c,s pronounciation audio dataset with all tones for my AI classification model project

Thumbnail
1 Upvotes