Synthetic Datasets

Explore our collection of synthetic datasets designed for learning and practicing data analysis, machine learning, and problem-solving skills. Each dataset comes with realistic scenarios and guided solutions.

Dataset Collection

19
Total Datasets
22
Topics Covered
3
Difficulty Levels

Available Datasets

Showing 19 of 19 datasets

Tackle blurry text, a broken archive entry, and fields that disappear in a scan.

5-8 hours
insuranceocrjsondata-validation

Start batch PDF extraction with short forms and CSV and JSON answers to check your work.

2-3 hours
insurancepdf-extractioncsvdata-validation

Practice OCR, keep leading zeroes in ZIP codes, and learn when a field simply is not on the page.

2-3 hours
shippingocrjsondata-validation

Build a batch job that can pick up where it left off, then check tracking numbers and addresses.

6-10 hours
shippingocrbatch-processingquality-assurance

The PDFs cover only part of the 14,956-transaction dataset. Finding that gap is part of the challenge.

6-10 hours
retailocrdata-cleaningreconciliation

Start with a receipt PDF, build a sales ledger, and uncover a pattern in loyalty discounts.

3-4 hours
retailpdf-extractioncsvreconciliation

Practice handling large files, unpacking nested items, and checking what the data actually covers.

5-8 hours
retailbig-datajsonreconciliation

Practice finding several receipts on one page and checking their item totals.

4-6 hours
retailocrjsonreconciliation

Six overlapping exports turn into one pull history and a chance to explore probability.

3-5 hours
jsondata-cleaningprobabilitystatistical-inference

Two PDFs and one model answer introduce field mapping, nested expenses, and simple checks.

1-2 hours
insurancepdf-extractionjsondata-validation

Practice OCR when a complete answer spreadsheet is not available.

4-6 hours
insuranceocrcsvquality-assurance

Use 100,000 sessions to study pageviews, visit times, and repeated trips to the Blog page.

3-4 hours
clickstreamcsvdata-analysissequence-analysis

Explore two ways to recover 15,347 receipts and fix a CSV with broken quoting.

5-8 hours
retailhtml-extractionocrreconciliation

A small import exercise where the file extension does not always tell you the format.

1-2 hours
transactionsdata-cleaningcsvreconciliation

Build a compact graph from 47,002 transactions and keep user labels separate from your review rules.

4-6 hours
transactionsgraph-analysiscsvdata-validation

Practice table extraction with shuffled text, gaps in IDs, and amounts in several currencies.

3-5 hours
transactionspdf-extractionreconciliationdata-validation

Catch duplicate files before they double your totals or sneak into both sides of a model evaluation.

2-4 hours
transactionsdeduplicationreconciliationdata-validation

A smaller payment network with tricky number formatting and useful lessons about risk labels.

3-5 hours
transactionspdf-extractiongraph-analysisreconciliation

Work through 2,392,287 transactions without mixing up store IDs or counting receipts from preview PDFs.

6-10 hours
retailbig-datacsvreconciliation