This guide will instruct you through:
- Creating your first R2 bucket and enabling its data catalog.
- Creating an API token needed for pipelines to authenticate with your data catalog.
- Creating your first pipeline with a simple ecommerce schema that writes to an Apache Iceberg ↗︎ table managed by Basin Catalog.
- Sending sample ecommerce data via HTTP endpoint.
- Validating data in your bucket and querying it with Basin SQL.
- Sign up for a Cloudflare account ↗︎.
- Install
Node.js↗︎.
Node.js version manager
Use a Node version manager like Volta ↗︎ or nvm ↗︎ to avoid permission issues and change Node.js versions. Wrangler, discussed later in this guide, requires a Node version of 16.17.0 or later.
-
If not already logged in, run:
npx wrangler loginyarn wrangler loginpnpm wrangler login -
Create an R2 bucket:
npx wrangler r2 bucket create pipelines-tutorialyarn wrangler r2 bucket create pipelines-tutorialpnpm wrangler r2 bucket create pipelines-tutorial
-
In the Cloudflare dashboard, go to the R2 object storage page.
Go to Overview ↗ -
Select Create bucket.
-
Enter the bucket name: pipelines-tutorial
-
Select Create bucket.
Enable the catalog on your R2 bucket:
npx wrangler basin catalog enable pipelines-tutorialyarn wrangler basin catalog enable pipelines-tutorialpnpm wrangler basin catalog enable pipelines-tutorialWhen you run this command, take note of the "Warehouse" and "Catalog URI". You will need these later.
-
In the Cloudflare dashboard, go to the R2 object storage page.
Go to Overview ↗ -
Select the bucket: pipelines-tutorial.
-
Switch to the Settings tab, scroll down to Basin Catalog, and select Enable.
-
Once enabled, note the Catalog URI and Warehouse name.
Basin Pipelines must authenticate to Basin Catalog with an R2 API token that has catalog and R2 permissions.
-
In the Cloudflare dashboard, go to the R2 object storage page.
Go to Overview ↗ -
Select Manage API tokens.
-
Select Create Account API token.
-
Give your API token a name.
-
Under Permissions, choose the Admin Read & Write permission.
-
Select Create Account API Token.
-
Note the Token value.
First, create a schema file that defines your ecommerce data structure:
Create schema.json:
{
"fields": [
{
"name": "user_id",
"type": "string",
"required": true
},
{
"name": "event_type",
"type": "string",
"required": true
},
{
"name": "product_id",
"type": "string",
"required": false
},
{
"name": "amount",
"type": "float64",
"required": false
}
]
}Use the interactive setup to create a pipeline that writes to Basin Catalog:
npx wrangler basin pipelines setupyarn wrangler basin pipelines setuppnpm wrangler basin pipelines setupFollow the prompts:
-
Pipeline name: Enter
ecommerce -
Stream configuration:
- Enable HTTP endpoint:
yes - Require authentication:
no(for simplicity) - Configure custom CORS origins:
no - Schema definition:
Load from file - Schema file path:
schema.json(or your file path)
- Enable HTTP endpoint:
-
Sink configuration:
- Destination type:
Data Catalog Table - R2 bucket name:
pipelines-tutorial - Namespace:
default - Table name:
ecommerce - Catalog API token: Enter your token from step 3
- Compression:
zstd - Roll file when size reaches (MB):
100 - Roll file when time reaches (seconds):
10(for faster data visibility in this tutorial)
- Destination type:
-
SQL transformation: Choose
Use simple ingestion queryto use:INSERT INTO ecommerce_sink SELECT * FROM ecommerce_stream
After setup completes, note the HTTP endpoint URL displayed in the final output.
-
In the Cloudflare dashboard, go to Basin Pipelines > Pipelines.
Go to Pipelines ↗ -
Select Create Pipeline.
-
Connect to a Stream:
- Pipeline name:
ecommerce - Enable HTTP endpoint for sending data: Enabled
- HTTP authentication: Disabled (default)
- Select Next
- Pipeline name:
-
Define Input Schema:
- Select JSON editor
- Copy in the schema:
{ "fields": [ { "name": "user_id", "type": "string", "required": true }, { "name": "event_type", "type": "string", "required": true }, { "name": "product_id", "type": "string", "required": false }, { "name": "amount", "type": "f64", "required": false } ] } - Select Next
-
Define Sink:
- Select your R2 bucket:
pipelines-tutorial - Storage type: Basin Catalog
- Namespace:
default - Table name:
ecommerce - Advanced Settings: Change Maximum Time Interval to
10 seconds - Select Next
- Select your R2 bucket:
-
Credentials:
- Disable Automatically create an Account API token for your sink
- Enter Catalog Token from step 3
- Select Next
-
Pipeline Definition:
- Leave the default SQL query:
INSERT INTO ecommerce_sink SELECT * FROM ecommerce_stream; - Select Create Pipeline
- Leave the default SQL query:
-
After pipeline creation, note the Stream ID for the next step.
Send ecommerce events to your pipeline's HTTP endpoint:
curl -X POST https://{stream-id}.ingest.cloudflare.com \
-H "Content-Type: application/json" \
-d '[
{
"user_id": "user_12345",
"event_type": "purchase",
"product_id": "widget-001",
"amount": 29.99
},
{
"user_id": "user_67890",
"event_type": "view_product",
"product_id": "widget-002"
},
{
"user_id": "user_12345",
"event_type": "add_to_cart",
"product_id": "widget-003",
"amount": 15.50
}
]'Replace {stream-id} with your actual stream endpoint from the pipeline setup.
-
In the Cloudflare dashboard, go to the R2 object storage page.
-
Select your bucket:
pipelines-tutorial. -
Confirm that your pipeline created Iceberg metadata and data files. If the files do not appear, wait a few minutes and try again.
-
The data is organized in the Apache Iceberg format with metadata tracking table versions.
Set up your environment to use Basin SQL:
export WRANGLER_BASIN_SQL_AUTH_TOKEN=YOUR_API_TOKENOr create a .env file with:
WRANGLER_BASIN_SQL_AUTH_TOKEN=YOUR_API_TOKENWhere YOUR_API_TOKEN is the token you created in step 3. For more information on setting environment variables, refer to Wrangler system environment variables.
Query your data:
npx wrangler basin sql query "YOUR_WAREHOUSE_NAME" "SELECT user_id, event_type, product_id, amount FROM default.ecommerce WHERE event_type = 'purchase' LIMIT 10"yarn wrangler basin sql query "YOUR_WAREHOUSE_NAME" "SELECT user_id, event_type, product_id, amount FROM default.ecommerce WHERE event_type = 'purchase' LIMIT 10"pnpm wrangler basin sql query "YOUR_WAREHOUSE_NAME" "SELECT user_id, event_type, product_id, amount FROM default.ecommerce WHERE event_type = 'purchase' LIMIT 10"Replace YOUR_WAREHOUSE_NAME with the warehouse name from step 2.
You can also query this table with any engine that supports Apache Iceberg. To connect other engines to Basin Catalog, refer to Connect to Iceberg engines.