Chatbot Dataset for Developers and Testers
This ready-made dataset contains 10,859 interactions across 3 tables, generated by GoMask DataFactory. It is suitable for developers and testers building or evaluating chatbot performance. A free sample is available, with full download costing credits.
At a glance
- 3tables
- 10,859rows
- 14columns
- Jan 2023 – Dec 2024date range
The 3 tables
preview and data dictionary per tableConversations conversations · table · 1,000 rows
Records individual turns within chatbot conversations.
Preview
| conversation_iduuid | user_iduuid | session_iduuid | timestampdatetime | user_utterancestring | chatbot_responsestring |
|---|---|---|---|---|---|
| e8b0639d-37d6-4321-a5c8-d306cb6713f9 | ab939f12-594e-4c4c-91c1-3b15bea6607e | 18f0e28a-121e-4db8-ab01-f8525ec9cfa1 | 2024-12-09 09:58:56 | Thanks, bye! | Goodbye! Have a great day! |
| 43299ef7-7e69-4117-8d4f-56f51848f508 | 6947aa55-0469-4233-ba37-703d1dcea5ae | 91cb57ef-32c4-4b89-be36-71dd90725cd5 | 2024-12-18 15:40:19 | See you later. | Farewell! Feel free to reach out anytime. |
| a311de76-7088-496b-a996-4955452fd186 | 6ad97635-448c-4b4d-afc0-f305cd2e584f | 67910d5e-9364-440d-9e49-701e7d127c42 | 2023-02-14 18:27:27 | Okay, thanks! | You're welcome! Have a good one. |
| 14da7003-a60a-48b8-b4da-77f76e639553 | fb65f8e7-b1d7-4a7e-b196-e47b6c7f61a1 | 514338be-f889-4b5b-8b03-2346aac8afe7 | 2023-03-29 07:43:05 | Hello, can you tell me what the current weather is like in London? | Certainly! The current weather in London is partly cloudy with a temperature of 15 degrees Celsius. There's a light breeze from the west. |
| 4cf5ccc5-4416-4702-877b-30d261dc1f80 | 11339a5a-6c53-4f1e-afb5-3d8807df7197 | 1ddf3939-fc3b-4d52-afc2-dabfd2e24ab8 | 2023-03-10 11:52:52 | Hi there! What's the capital of France? | Hello! The capital of France is Paris. Is there anything else I can help you with today? |
| be781499-dc24-4967-bb82-29182d6f9614 | 1b5b6b4a-e8fb-4156-9812-491f1d7199e6 | ba7e7a16-c4d5-47d6-87ab-379b3b7d0af7 | 2024-11-08 10:12:14 | Good morning. Could you please provide the latest stock price for Apple? | Good morning! The latest stock price for Apple (AAPL) is $175.50. Please note that stock prices can fluctuate. |
| 8c2a0bbd-e1b0-4076-93c2-d6fc5d0efe2e | 23edb083-e260-45e6-925c-f0bcfaeeaf4e | e9e830d6-64fe-4403-93fc-cbb942699688 | 2024-10-18 11:35:17 | Excuse me, what time is it in Tokyo right now? | Certainly. The current time in Tokyo, Japan is 11:35 PM on October 18, 2024. It's quite late there! |
| be465d2b-5c98-41b5-97ea-db454f448e24 | 034e44c7-3c9d-403c-a88e-9cd157bf2c1f | 3d318ca5-a7a8-4589-a0af-1f54e5f166a0 | 2024-04-17 10:18:41 | hi | Hello! How can I help you today? |
| a7c98c5f-1096-4c0d-9b46-ed238c6b56f8 | cdf0bc83-9349-4982-99c0-4e4dcdce2b57 | 79f143e6-91bf-4771-bead-aa7452eb75bb | 2023-08-01 14:52:52 | ok | Great! Is there anything else I can assist you with? |
| 12a7c840-d017-4f3f-8231-fc6c33f173df | 803e5bb8-dbc6-46ef-874e-343f88bd8431 | 0980865c-401d-4cc9-8bd2-b7d38ecc2824 | 2024-10-26 12:48:43 | help | I'm here to help. What do you need assistance with? |
Data dictionary
| column | type | description | example | null % |
|---|---|---|---|---|
conversation_ | uuid | Unique identifier for each conversation turn.unique | e8b0639d-37d6-4321-a5c8-d306cb6713f9 | 0% |
user_ | uuid | Identifier of the user participating in the chat session.unique | ab939f12-594e-4c4c-91c1-3b15bea6607e | 0% |
session_ | uuid | Unique identifier for the continuous dialogue session.unique | 18f0e28a-121e-4db8-ab01-f8525ec9cfa1 | 0% |
timestamp | datetime | Timestamp when the interaction turn took place.unique | 2024-12-09 09:58:56 | 0% |
user_ | string | Natural language text input submitted by the user. | Thanks, bye! | 0% |
chatbot_ | string | Natural language response synthesized or returned by the chatbot. | Goodbye! Have a great day! | 0% |
Intents intents · fact table · 4,894 rows
Details the intents detected from user utterances and their confidence scores.
Preview
| intent_iduuid | detected_intentstring | confidence_scoredecimal | conversation_iduuid |
|---|---|---|---|
| d2aa77fb-867b-48a1-8383-d1102e9ae98c | request_info | 0.86 | e8b0639d-37d6-4321-a5c8-d306cb6713f9 |
| 2806c7cc-1aa3-4289-9a0d-11bda5bd69cf | request_info | 0.94 | 8c2a0bbd-e1b0-4076-93c2-d6fc5d0efe2e |
| d81ce89a-5c65-45b7-9774-cf296914210b | request_info | 0.85 | 12a7c840-d017-4f3f-8231-fc6c33f173df |
| 7cbcfa08-82d9-4b3c-8ccd-731836e51e12 | request_info | 0.93 | 12a7c840-d017-4f3f-8231-fc6c33f173df |
| 5724d424-150d-4b1e-8ea4-1b42c96f3661 | request_info | 0.9 | a7c98c5f-1096-4c0d-9b46-ed238c6b56f8 |
| 8a9532de-25a9-406c-9e34-71fb8bf8caa0 | request_info | 0.85 | 14da7003-a60a-48b8-b4da-77f76e639553 |
Data dictionary
| column | type | description | example | null % |
|---|---|---|---|---|
intent_ | uuid | Primary key uniquely identifying the detected intent record.unique | 2290a498-7df0-4025-8938-4a84686c381c | 0% |
detected_ | string | The specific intent category detected from the user's conversational turn. | request_info | 0% |
confidence_ | decimal | Natural language understanding model confidence score for the detected intent. | 0.95 | 0% |
conversation_ | uuid | Foreign key to conversations.conversation_id. | ec5c3b20-95f6-4d73-993b-d28ee73bab46 | 0% |
Responses responses · fact table · 4,965 rows
Categorizes the types of responses provided by the chatbot.
Preview
| response_iduuid | response_typestring | response_textstring | conversation_iduuid |
|---|---|---|---|
| 033dfd50-8c4e-4f30-90c4-38271545e6cf | standard | Completed. The task has been successfully executed. | be781499-dc24-4967-bb82-29182d6f9614 |
| bf88de11-0c46-410f-83fa-1187021f5aeb | standard | Good morning! I'm here and ready to assist you. What's on your mind? | 43299ef7-7e69-4117-8d4f-56f51848f508 |
| 87379ba9-04bd-4317-8623-70c42e05a0a4 | standard | Thank you for providing the necessary details. I can now proceed. | 8c2a0bbd-e1b0-4076-93c2-d6fc5d0efe2e |
| 6dab7115-6131-4250-afe5-7f9ef35fa436 | standard | The process has been initiated as per your instructions. We'll keep you updated. | a7c98c5f-1096-4c0d-9b46-ed238c6b56f8 |
Data dictionary
| column | type | description | example | null % |
|---|---|---|---|---|
response_ | uuid | Unique primary key identifier for the specific chatbot response segment.unique | fa66ee2f-65b2-4d5b-a107-a4e33adb06b6 | 0% |
response_ | string | Type of chatbot response generated: standard, clarification, error, or follow_up. | standard | 0% |
response_ | string | The natural language text segment returned by the chatbot to the user. | Hello! I can help you with your inquiries today. How may I assist you? | 0% |
conversation_ | uuid | Foreign key to conversations.conversation_id. | 81e07aee-9101-4219-970a-43f44c4eeeb3 | 0% |
How the tables join
intents.conversation_id references conversations.conversation_idmany to one: each Intents row points to one Conversations rowresponses.conversation_id references conversations.conversation_idmany to one: each Responses row points to one Conversations row
Questions to answer with it
Analyze the distribution of detected intents across all conversations. Which intents are most common?
Join conversations and intents tables, then group by detected_intent and count occurrences.
tables: conversations, intents
Calculate the average confidence score for each detected intent. Which intents have the highest and lowest average confidence?
Group intents by detected_intent and calculate the average of confidence_score.
tables: intents
Identify conversations where the chatbot response type is 'error' and analyze the corresponding user utterances.
Join conversations and responses tables, filter for response_type = 'error', and select user_utterance.
tables: conversations, responses
For a specific user_id, retrieve their conversation history, including user utterances and chatbot responses, ordered by timestamp.
Filter the conversations table by a specific user_id and order by timestamp.
tables: conversations
Starter SQL
run against this data before publishingTable names match the SQLite file and the SQL script.
Top 5 Most Common Intents
SELECT detected_intent, COUNT(*) AS intent_count FROM intents GROUP BY detected_intent ORDER BY intent_count DESC LIMIT 5;Average Confidence Score per Intent
SELECT detected_intent, AVG(confidence_score) AS average_confidence FROM intents GROUP BY detected_intent ORDER BY average_confidence DESC LIMIT 20;Error Responses and User Utterances
SELECT c.user_utterance, r.response_text FROM conversations c JOIN responses r ON c.conversation_id = r.conversation_id WHERE r.response_type = 'error' LIMIT 20;User Conversation History
SELECT conversation_id, timestamp, user_utterance, chatbot_response FROM conversations WHERE user_id = 'ab939f12-594e-4c4c-91c1-3b15bea6607e' ORDER BY timestamp LIMIT 20;Intent Distribution with Window Function
SELECT detected_intent, COUNT(*) AS intent_count, RANK() OVER (ORDER BY COUNT(*) DESC) as rank FROM intents GROUP BY detected_intent LIMIT 20;Load it with pandas
import pandas as pd
# Unzip the CSV download first: one file per table
conversations = pd.read_csv("conversations.csv")
intents = pd.read_csv("intents.csv")
responses = pd.read_csv("responses.csv")
# Join intents to conversations
df = intents.merge(conversations, left_on="conversation_id", right_on="conversation_id", how="left", suffixes=("", "_conversations"))
print(df.groupby("conversation_id").size().describe())Using it in your tool
- Excel
- Load each table into a separate Excel sheet. Use XLOOKUP to join tables based on conversation_id. Create a PivotTable on the 'intents' sheet to analyze intent distribution and average confidence scores.
- Power BI
- Load all three tables. Create relationships: conversations.conversation_id to intents.conversation_id and conversations.conversation_id to responses.conversation_id. Create measures for total conversations, average confidence score, and count of error responses.
- SQL
- Load the SQLite or SQL dump into your preferred database. Use the provided table and column names for queries. Joins are typically performed on conversation_id. Primary keys are listed in the table schemas.
- Python
- Use pandas to load CSV or Parquet files. Merge tables using conversation_id. Analyze intent distributions, confidence scores, and response types using DataFrame operations.
Formats available
- CSV (zip)One CSV file per table, zipped
- Excel workbookOne worksheet per table
- SQLite databaseA ready-to-query database file with every table
- Parquet (zip)One Parquet file per table, zipped
- SQL scriptCREATE TABLE with primary and foreign keys, then INSERTs
- CSVA single CSV file
- JSONA single JSON file
How this data was generated
Synthetic data. Every row was generated: no real people, customers or companies are in this dataset.
Synthetic data generated by GoMask DataFactory from a relational blueprint: keys, links and rules are enforced in code, text columns are filled by a language model. No real people, companies or transactions.
- Data generated using GoMask DataFactory.
- Synthetic data mimics chatbot interaction patterns.
- Includes conversation turns, detected intents, and chatbot responses.
- Checked by an automated quality gate: unique keys, no orphan foreign keys, required columns filled, declared rules and date ranges (realism score 92).
Limitations
- The data is synthetic and does not represent real-world user conversations.
- Does not include user profiles or detailed session metadata beyond session_id.
- Limited to 8 distinct intents and 4 response types.
- Distributions and correlations are modelled, not measured from real records.
blueprint · chatbot-dataset
Scale this dataset
Same tables. As many rows as you need.
Open the blueprint behind these 3 tables in Data Factory: keep the relationships, change a column, and generate it at the size you need.
- 50,000 rows
- 200,000 rows
- 1,000,000 rows
- Tables
- conversations, intents, responses
- Licence
- yours to use, including commercially
- API slug
- chatbot-dataset