Skip to main content
This dataset contains over 150M customer reviews of Amazon products. The data is in snappy-compressed Parquet files in AWS S3 that total 49GB in size (compressed).
You can try this on self-hosted ClickHouse or on ClickHouse Cloud. If you don’t already have an account, you can start a free ClickHouse Cloud trial.
Let’s walk through the steps to insert the data into ClickHouse.

Loading the dataset

  1. Without inserting the data into ClickHouse, we can query it in place. Let’s grab some rows, so we can see what they look like:
The rows look like:
  1. Let’s define a new MergeTree table named amazon_reviews to store this data in ClickHouse:
  1. The following INSERT command uses the s3Cluster table function, which distributes the S3 files among the nodes of your cluster so they are read in parallel.
In ClickHouse Cloud, your cluster is named default and you can run this as shown. On a self-managed server, replace default with the name of your cluster — or, if you are running a single server, use the s3 table function instead, dropping the first argument so the call starts with the URL. We use a wildcard in the URL to insert all matching review files.
  1. That query doesn’t take long - averaging about 300,000 rows per second. Within 5 minutes or so you should see all the rows inserted:
  1. Let’s see how much space our data is using:
The original data was about 70G, but compressed in ClickHouse it takes up about 30G.

Example queries

  1. Let’s run some queries. Here are the top 10 most-helpful reviews in the dataset:
This query is using a projection to speed up performance.
  1. Here are the top 10 products in Amazon with the most reviews:
  1. Here are the average review ratings per month for each product (an actual Amazon job interview question!):
  1. Here are the total number of votes per product category. This query is fast because product_category is in the primary key:
  1. Let’s find the products with the word “awful” occurring most frequently in the review. This is a big task - over 151M strings have to be parsed looking for a single word:
runnable
Notice the query time for such a large amount of data. The results are also a fun read!
  1. We can run the same query again, except this time we search for awesome in the reviews:
runnable
Last modified on September 25, 2026