Glossary

The words, defined once.

111 terms across test data, synthetic data and compliance. Each has a short definition here and a full page behind it.

A

API Test Data

Test data specifically designed for API testing, including request payloads, response data, edge cases, and error scenarios to validate API functionality and contracts.

Full definition →

C

CI/CD Integration

Continuous Integration and Continuous Deployment practices that automate software building, testing, and deployment, enabling rapid and reliable delivery of applications.

Full definition →

CI/CD Test Data

Test data provisioned automatically as part of Continuous Integration and Continuous Deployment pipelines, enabling automated testing with fresh, compliant data.

Full definition →

Cloud Test Data

Test data management practices optimized for cloud environments, leveraging cloud storage, compute elasticity, and infrastructure-as-code for scalable test data operations.

Full definition →

D

Data Anonymization

The process of permanently removing or modifying personally identifiable information from datasets so that individuals cannot be re-identified, even with additional data sources.

Full definition →

Data Integrity

The accuracy, consistency, and reliability of data throughout its lifecycle, ensuring data remains correct, complete, and trustworthy for its intended use.

Full definition →

Data Lake

A centralized repository that stores vast amounts of raw, unstructured, and structured data in its native format until needed for analysis.

Full definition →

Data Masking

The process of replacing sensitive data with realistic but fictional alternatives to protect privacy while maintaining data usability for testing and development.

Full definition →

Data Migration

The process of moving data from one system, storage, format, or database to another, typically during technology upgrades, consolidations, or cloud transitions.

Full definition →

Data Pipeline

An automated workflow that moves data from source systems through various processing stages to destination systems, often including extraction, transformation, validation, and loading steps.

Full definition →

Data Quality

The measure of data's fitness for its intended purpose, evaluated across dimensions like accuracy, completeness, consistency, timeliness, and validity.

Full definition →

Data Redaction

The process of removing or obscuring sensitive portions of data, typically replacing them with blocks, asterisks, or complete removal of the information.

Full definition →

Data Residency

Legal and compliance requirements that data must be stored and processed within specific geographic boundaries, affecting test data location and movement.

Full definition →

Data Scrambling

A data masking technique that rearranges or shuffles data values within a column to protect privacy while maintaining data distribution and statistical properties.

Full definition →

Data Subsetting

The process of extracting a smaller, representative portion of a production database while maintaining referential integrity and business logic relationships.

Full definition →

Data Substitution

A masking technique that replaces real sensitive values with realistic fake alternatives from predefined lists or generated data, maintaining data format and usability.

Full definition →

Data Warehouse

A centralized repository optimized for analytical queries and reporting, integrating data from multiple sources for business intelligence and decision-making.

Full definition →

Database Schema

The logical structure and organization of a database including tables, columns, data types, relationships, constraints, and indexes that define how data is stored and accessed.

Full definition →

De-identification

The process of removing or obscuring personal identifiers from data to prevent identification of individuals, while retaining data utility for analysis and testing.

Full definition →

E

ETL Testing

Validation of Extract, Transform, Load processes that move and transform data between systems, requiring test data representing source formats, transformation rules, and target schemas.

Full definition →

F

Foreign Key

A database constraint that creates a link between two tables by requiring values in one table to match values in a primary key column of another table.

Full definition →

G

H

I

L

M

Mock Data

Simple, minimal fake data used primarily in unit testing and development, prioritizing speed and simplicity over realism or statistical accuracy.

Full definition →

N

NoSQL Databases

Non-relational databases designed for scalability, flexibility, and performance with unstructured or semi-structured data, using models like document, key-value, column-family, or graph.

Full definition →

P

Primary Key

A column or combination of columns in a database table that uniquely identifies each row, serving as the main reference point for relationships with other tables.

Full definition →

Privacy by Design

An approach to systems engineering that takes privacy into account throughout the entire development process, embedding privacy protections into technology from the outset.

Full definition →

Production-like Data

Test data that accurately mimics production data characteristics including formats, distributions, relationships, and business rules without containing actual sensitive information.

Full definition →

Pseudonymization

A data protection technique that replaces identifying information with artificial identifiers (pseudonyms), allowing data subjects to be indirectly identified only with additional information kept separately.

Full definition →

R

S

Sensitive Data

Information that must be protected from unauthorized access due to legal, ethical, or business requirements, including personally identifiable information, financial data, health records, and proprietary business information.

Full definition →

SOX (Sarbanes-Oxley Act)

A United States federal law enacted in 2002 that mandates strict financial record-keeping and reporting requirements for public companies to protect investors from fraudulent financial practices.

Full definition →

T

Test Data Generation

The process of creating test data from scratch using algorithms, patterns, and rules rather than copying from production, encompassing synthetic data, mock data, and procedural generation.

Full definition →

Test Data Strategy

A comprehensive plan defining how an organization will acquire, protect, manage, and provision test data to support testing activities while meeting compliance and operational requirements.

Full definition →

Tokenization

A security technique that replaces sensitive data with non-sensitive substitutes (tokens) that have no exploitable value, with the original data stored securely in a separate token vault.

Full definition →

U

Unique Constraints

Database rules that ensure a column or combination of columns contains only unique values across all rows, preventing duplicates in fields like email addresses or usernames.

Full definition →

Know the words. Now make the rows.

Every term above shows up in a real run.