Datasets by use case

Classification Datasets for Machine Learning

Each dataset here has a column to predict: is_fraud, churned, attrition, readmitted, default_flag. The features around it are the ones a real model would use, and the sample rows on each page show you the class balance before you download. Load a CSV into pandas or scikit-learn and start with a baseline.

Fraud

Card and payment transactions with an is_fraud or fraud_label column, plus the features fraud models use: amount, merchant category, entry mode, international flag, device and velocity counts. Judge models on precision and recall, not accuracy.

Finance17 cols · 200 rows

Credit Card Fraud Detection

This dataset provides detailed, labeled records of simulated credit card transactions, including transaction amounts, merchant and cardholder information, and fraud indicators. It is ideal for developing and benchmarking machine learning models aimed at detecting fraudulent activity and reducing financial risk in payment systems. The inclusion of transaction context and cardholder demographics supports advanced analytics and feature engineering.

merchant_categorytransaction_typeentry_mode
grocerypurchasechip
coffee_shoppurchasecontactless
electronicspurchaseonline
transaction_idcard_numbertransaction_datetime+11 more
Finance19 cols · 200 rows

Financial Transaction Fraud Features

This dataset provides a detailed, feature-rich record of synthetic banking transactions, including transaction metadata, account and merchant information, contextual behavioral features, and fraud labels. It is ideal for developing, training, and benchmarking machine learning models for fraud detection and anomaly analysis in financial services.

transaction_typemerchant_categoryfraud_type
purchasegroceries–
refundelectronicsfake_merchant
purchaseclothing–
transaction_idaccount_idtransaction_datetime+13 more
Finance19 cols · 200 rows

Payment Fraud Detection Dataset

This dataset contains detailed synthetic payment transaction records, each labeled with ground-truth indicators of fraud. It includes transaction metadata, customer and merchant identifiers, payment methods, device and location context, and fraud reasons, making it ideal for developing and benchmarking machine learning models for payment fraud detection and risk mitigation.

currencytransaction_statusfraud_label
USDcompletedfalse
GBPcompletedfalse
INRfailedtrue
transaction_idtransaction_datetimeamount+13 more
Insurance20 cols · 200 rows

Medical Insurance Fraud Detection

This dataset contains detailed synthetic records of medical insurance claims, including patient demographics, provider information, claim amounts, service dates, and labeled indicators of fraudulent activity. Designed for machine learning and analytics, it enables robust research and development of fraud detection models in healthcare and insurance. The dataset supports granular analysis of claim patterns, provider behaviors, and patient demographics to identify and prevent fraudulent claims.

claim_statuspatient_genderprovider_specialty
approvedfemaleFamily Practice
approvedmalePsychiatry
approvedfemaleCardiology
claim_idpatient_idprovider_id+14 more

Churn and attrition

Customers and employees who left, with the columns that hint why: tenure, satisfaction, usage, pay and policy details. Try logistic regression first, then a gradient-boosted model, and compare what each says about the drivers.

Human Resources26 cols · 200 rows

Employee Attrition Prediction Dataset

This dataset provides detailed HR records for employees, including demographics, job history, satisfaction, performance, and attrition status. It is ideal for building predictive models to identify turnover risks, analyze workforce trends, and inform retention strategies. Rich features enable advanced analytics for HR decision-making and organizational planning.

first_namelast_namemarital_status
RiyaPatelMarried
CarlosDominguezSingle
AishaOkaforWidowed
employee_idgenderdate_of_birth+20 more
Insurance31 cols · 200 rows

Insurance Policyholder Churn Insights

This dataset provides a comprehensive view of insurance policyholders, their demographic details, policy information, claims history, and churn status for both life and auto insurance products. It is designed to support predictive modeling of customer attrition, enabling insurers to identify at-risk customers and develop targeted retention strategies. The inclusion of satisfaction scores, contact history, and churn reasons makes it ideal for advanced analytics and customer experience optimization.

first_namelast_namepolicy_type
AmeliaRichardsauto
OmarAl-Mansourilife
PriyaMehraauto
policyholder_iddate_of_birthgender+25 more
Gaming16 cols · 200 rows

Game Session Telemetry Dataset

This dataset contains detailed logs of simulated online gaming sessions, including player identifiers, session timings, actions, feature usage, purchases, and outcomes. Designed for gaming analytics, it enables churn prediction, player segmentation, and feature adoption analysis, offering valuable insights for game developers and startups.

player_segmentsession_outcomedevice_type
newabandonedmobile
casualwinmobile
newabandonedconsole
session_idplayer_idgame_id+10 more

Risk: readmission and default

Hospital admissions with a readmitted flag, and loans with repayment histories and default flags. Both reward careful handling of dates, so nothing from the future leaks into your features.

Healthcare22 cols · 200 rows

Hospital Readmission Prediction

This dataset provides detailed, labeled records of hospital admissions, including patient demographics, diagnoses, procedures, and outcomes related to readmission. It is designed to support predictive modeling for hospital readmissions, enabling health systems to identify risk factors, optimize care transitions, and reduce costs. The comprehensive schema facilitates advanced analytics and machine learning applications in healthcare quality improvement.

readmittedadmission_typeinsurance_type
falseelectiveprivate
trueemergencymedicare
falseemergencymedicare
patient_idadmission_idadmission_date+16 more
Finance30 cols · 200 rows

Synthetic Bank Loan Repayment Patterns

This dataset provides detailed, anonymized records of bank loan origination, repayment, default, and prepayment patterns, enriched with borrower demographic and employment information. It is ideal for developing and validating credit scoring models, analyzing risk factors, and exploring repayment behaviors across diverse borrower segments. The data structure supports advanced analytics and machine learning applications in financial risk management.

repayment_statusborrower_marital_statusborrower_employment_status
currentsingleemployed_full_time
paid_offsingleemployed_part_time
currentmarriedemployed_full_time
loan_idborrower_idorigination_date+24 more
Finance33 cols · 200 rows

Finance Peer-to-Peer Lending Data

This dataset provides detailed, transaction-level records from a peer-to-peer lending platform, including loan terms, borrower and investor attributes, credit scores, repayment history, and loan performance indicators. It is ideal for credit risk modeling, investor analytics, and financial platform optimization, supporting both operational and research applications in alternative lending.

loan_statusborrower_employment_statusinvestor_type
activeemployedindividual
completedretiredinstitutional
activestudentindividual
loan_idborrower_idinvestor_id+27 more

Every dataset in this collection

10 datasets

Finance17 cols · 200 rows

Credit Card Fraud Detection

This dataset provides detailed, labeled records of simulated credit card transactions, including transaction amounts, merchant and cardholder information, and fraud indicators. It is ideal for developing and benchmarking machine learning models aimed at detecting fraudulent activity and reducing financial risk in payment systems. The inclusion of transaction context and cardholder demographics supports advanced analytics and feature engineering.

merchant_categorytransaction_typeentry_mode
grocerypurchasechip
coffee_shoppurchasecontactless
electronicspurchaseonline
transaction_idcard_numbertransaction_datetime+11 more
Finance19 cols · 200 rows

Financial Transaction Fraud Features

This dataset provides a detailed, feature-rich record of synthetic banking transactions, including transaction metadata, account and merchant information, contextual behavioral features, and fraud labels. It is ideal for developing, training, and benchmarking machine learning models for fraud detection and anomaly analysis in financial services.

transaction_typemerchant_categoryfraud_type
purchasegroceries–
refundelectronicsfake_merchant
purchaseclothing–
transaction_idaccount_idtransaction_datetime+13 more
Finance19 cols · 200 rows

Payment Fraud Detection Dataset

This dataset contains detailed synthetic payment transaction records, each labeled with ground-truth indicators of fraud. It includes transaction metadata, customer and merchant identifiers, payment methods, device and location context, and fraud reasons, making it ideal for developing and benchmarking machine learning models for payment fraud detection and risk mitigation.

currencytransaction_statusfraud_label
USDcompletedfalse
GBPcompletedfalse
INRfailedtrue
transaction_idtransaction_datetimeamount+13 more
Insurance20 cols · 200 rows

Medical Insurance Fraud Detection

This dataset contains detailed synthetic records of medical insurance claims, including patient demographics, provider information, claim amounts, service dates, and labeled indicators of fraudulent activity. Designed for machine learning and analytics, it enables robust research and development of fraud detection models in healthcare and insurance. The dataset supports granular analysis of claim patterns, provider behaviors, and patient demographics to identify and prevent fraudulent claims.

claim_statuspatient_genderprovider_specialty
approvedfemaleFamily Practice
approvedmalePsychiatry
approvedfemaleCardiology
claim_idpatient_idprovider_id+14 more
Human Resources26 cols · 200 rows

Employee Attrition Prediction Dataset

This dataset provides detailed HR records for employees, including demographics, job history, satisfaction, performance, and attrition status. It is ideal for building predictive models to identify turnover risks, analyze workforce trends, and inform retention strategies. Rich features enable advanced analytics for HR decision-making and organizational planning.

first_namelast_namemarital_status
RiyaPatelMarried
CarlosDominguezSingle
AishaOkaforWidowed
employee_idgenderdate_of_birth+20 more
Insurance31 cols · 200 rows

Insurance Policyholder Churn Insights

This dataset provides a comprehensive view of insurance policyholders, their demographic details, policy information, claims history, and churn status for both life and auto insurance products. It is designed to support predictive modeling of customer attrition, enabling insurers to identify at-risk customers and develop targeted retention strategies. The inclusion of satisfaction scores, contact history, and churn reasons makes it ideal for advanced analytics and customer experience optimization.

first_namelast_namepolicy_type
AmeliaRichardsauto
OmarAl-Mansourilife
PriyaMehraauto
policyholder_iddate_of_birthgender+25 more
Gaming16 cols · 200 rows

Game Session Telemetry Dataset

This dataset contains detailed logs of simulated online gaming sessions, including player identifiers, session timings, actions, feature usage, purchases, and outcomes. Designed for gaming analytics, it enables churn prediction, player segmentation, and feature adoption analysis, offering valuable insights for game developers and startups.

player_segmentsession_outcomedevice_type
newabandonedmobile
casualwinmobile
newabandonedconsole
session_idplayer_idgame_id+10 more
Healthcare22 cols · 200 rows

Hospital Readmission Prediction

This dataset provides detailed, labeled records of hospital admissions, including patient demographics, diagnoses, procedures, and outcomes related to readmission. It is designed to support predictive modeling for hospital readmissions, enabling health systems to identify risk factors, optimize care transitions, and reduce costs. The comprehensive schema facilitates advanced analytics and machine learning applications in healthcare quality improvement.

readmittedadmission_typeinsurance_type
falseelectiveprivate
trueemergencymedicare
falseemergencymedicare
patient_idadmission_idadmission_date+16 more
Finance30 cols · 200 rows

Synthetic Bank Loan Repayment Patterns

This dataset provides detailed, anonymized records of bank loan origination, repayment, default, and prepayment patterns, enriched with borrower demographic and employment information. It is ideal for developing and validating credit scoring models, analyzing risk factors, and exploring repayment behaviors across diverse borrower segments. The data structure supports advanced analytics and machine learning applications in financial risk management.

repayment_statusborrower_marital_statusborrower_employment_status
currentsingleemployed_full_time
paid_offsingleemployed_part_time
currentmarriedemployed_full_time
loan_idborrower_idorigination_date+24 more
Finance33 cols · 200 rows

Finance Peer-to-Peer Lending Data

This dataset provides detailed, transaction-level records from a peer-to-peer lending platform, including loan terms, borrower and investor attributes, credit scores, repayment history, and loan performance indicators. It is ideal for credit risk modeling, investor analytics, and financial platform optimization, supporting both operational and research applications in alternative lending.

loan_statusborrower_employment_statusinvestor_type
activeemployedindividual
completedretiredinstitutional
activestudentindividual
loan_idborrower_idinvestor_id+27 more

Questions

How many rows does each dataset have?
A few hundred, enough to build and debug a pipeline. For training at scale, open the dataset in GoMask Data Factory and generate up to a million rows with the same columns, at the class balance you ask for.
Is the data real?
No. It is synthetic data generated to behave like the real thing, so it is safe to share, publish and use in teaching.
Which format should I use in Python?
CSV or Parquet both load with pandas in one line. Excel, JSON, JSONL, SQL, TSV and XML are available too.

What should your data show?

Preview 20 rows free
No signup. No card.