Back to Glossary
🔄Synthetic Data Generation

Data Subsetting

Quick Definition

The process of extracting a smaller, representative portion of a production database while maintaining referential integrity and business logic relationships.

What is Data Subsetting?

Data Subsetting creates reduced-size copies of production databases by selecting specific rows based on criteria while preserving the relationships between tables. Instead of copying entire multi-terabyte production databases to test environments, subsetting extracts only the data needed for testing - dramatically reducing storage costs, improving performance, and accelerating provisioning.

Intelligent subsetting algorithms start with seed criteria (e.g., "customers from California created in the last year") and then traverse foreign key relationships to include all related records across dependent tables. This ensures the subset remains logically consistent: if you include an order, the subsetting process automatically includes the related customer, products, shipping address, and payment records.

Subsetting strategies include: Horizontal Subsetting (selecting specific rows, e.g., 10% of customers), Vertical Subsetting (selecting specific columns), Referential Subsetting (following foreign key relationships), and Conditional Subsetting (applying business rules like geography or time periods). Advanced subsetting maintains statistical representativeness, ensuring the subset accurately reflects production data distributions.

Organizations use subsetting to solve multiple challenges: reducing 5TB production databases to 50GB test databases (90%+ storage savings), accelerating database provisioning from hours to minutes, enabling laptop-based development with realistic data, lowering cloud infrastructure costs, and improving test execution speed. Subsetting is often combined with data masking to protect the extracted data.

Common Use Cases

  • Large database test environment provisioning
  • Developer local database creation
  • Performance testing with manageable datasets
  • Cloud cost optimization
  • Offshore development team data delivery

🎯How GoMask Helps

GoMask provides intelligent data subsetting that automatically follows foreign key relationships to maintain referential integrity. Define subsetting rules once and our engine extracts logically consistent data subsets. Combine subsetting with masking to deliver small, fast, compliant test databases.

Need help with Data Subsetting?

GoMask makes realistic synthetic datasets with the patterns you ask for. Get started in minutes.