Clean Apify scraper results into a CSV
A scraper's raw results usually need the same clean-up before they go into a spreadsheet or CRM: drop the rows you don't want, remove repeats, pick and rename the columns. This page walks through one example with my Apify Actor, Dataset Transformer, from the raw rows to the CSV.
Dataset Transformer on the Apify Store
The scraper's results
A Google Maps-style scraper ran two searches over the same area, "coffee shop in Austin, TX" and "cafe in Austin, TX". Places that match both searches come back twice, and each row has 24 fields, most of them nested or not needed in a spreadsheet: opening hours, image lists, the search rank and the scrape time. The places, phone numbers (555-01xx) and websites (.example) are made up for this example.
The goal: one CSV with one row per open place rated 4.3 or higher that has a phone number, the most reviewed first, with plain column names.
Before: 15 rows
Six of the 24 fields, and what happens to each row:
| # | title | What happens | totalScore | reviewsCount | phone | permanentlyClosed | searchString |
|---|---|---|---|---|---|---|---|
| 1 | Juniper & Oak Coffee | Kept | 4.7 | 1284 | (512) 555-0101 | false | coffee shop in Austin, TX |
| 2 | Lantern Street Roasters | Kept | 4.6 | 902 | (512) 555-0102 | false | coffee shop in Austin, TX |
| 3 | Copper Kettle Café | Kept (no website: empty cell) | 4.4 | 377 | (512) 555-0103 | false | coffee shop in Austin, TX |
| 4 | Night Owl Espresso Bar | Filtered out: rated below 4.3 | 4.1 | 215 | (512) 555-0104 | false | coffee shop in Austin, TX |
| 5 | Riverbend Bakery & Coffee | Kept | 4.8 | 2140 | (512) 555-0105 | false | coffee shop in Austin, TX |
| 6 | The Grind House | Filtered out: permanently closed | 4.5 | 640 | (512) 555-0106 | true | coffee shop in Austin, TX |
| 7 | Bluebonnet Brew Co. | Filtered out: no phone | 4.5 | 88 | null | false | coffee shop in Austin, TX |
| 8 | Sixth & Vine Coffee | Kept | 4.5 | 312 | (512) 555-0108 | false | coffee shop in Austin, TX |
| 9 | Arbor Lane Café | Kept | 4.6 | 59 | (512) 555-0109 | false | coffee shop in Austin, TX |
| 10 | Riverbend Bakery & Coffee | Duplicate of 5 (same placeId) | 4.8 | 2140 | (512) 555-0105 | false | cafe in Austin, TX |
| 11 | Mesquite Moon Café | Kept | 4.3 | 141 | (512) 555-0110 | false | cafe in Austin, TX |
| 12 | Juniper & Oak Coffee | Duplicate of 1 | 4.7 | 1284 | (512) 555-0101 | false | cafe in Austin, TX |
| 13 | Pecan Tree Tea & Coffee | Filtered out: rated below 4.3 | 4.2 | 403 | (512) 555-0111 | false | cafe in Austin, TX |
| 14 | Lantern Street Roasters | Duplicate of 2 | 4.6 | 902 | (512) 555-0102 | false | cafe in Austin, TX |
| 15 | Copper Kettle Café | Duplicate of 3 | 4.4 | 377 | (512) 555-0103 | false | cafe in Austin, TX |
The input
Three filters, duplicates removed by the place's ID, sorted by review count, and twelve columns kept, eight of them renamed (old -> new). Paste it into the Actor's JSON input, or fill in the same values in the form, with your scraper run's dataset in Apify dataset:
{
"datasetId": "YOUR-SCRAPER-DATASET-ID",
"filters": ["permanentlyClosed != true", "totalScore >= 4.3", "phone is not empty"],
"dedupeBy": ["placeId"],
"sortBy": ["reviewsCount desc"],
"fields": ["title -> name", "categoryName -> category", "totalScore -> rating", "reviewsCount -> reviews",
"phone", "website", "street", "city", "postalCode -> zip",
"location.lat -> lat", "location.lng -> lng", "url -> googleMapsUrl"],
"exportFormat": "csv"
}
Filters, duplicates and sorting use the scraper's own field names (totalScore, placeId). The new names only appear in the output. location.lat reaches into the nested location object.
After: 7 rows
The run's status line: "7 rows written from 15 rows read; 4 filtered out; 4 duplicates removed; file saved as OUTPUT.csv". Six of the twelve columns (the others are category, street, city, lat, lng and googleMapsUrl):
name | rating | reviews | phone | website | zip |
|---|---|---|---|---|---|
| Riverbend Bakery & Coffee | 4.8 | 2140 | (512) 555-0105 | https://riverbendbakery.example | 78746 |
| Juniper & Oak Coffee | 4.7 | 1284 | (512) 555-0101 | https://juniperoak.example | 78702 |
| Lantern Street Roasters | 4.6 | 902 | (512) 555-0102 | https://lanternstreet.example/menu?utm_source=gmb | 78704 |
| Copper Kettle Café | 4.4 | 377 | (512) 555-0103 | 78741 | |
| Sixth & Vine Coffee | 4.5 | 312 | (512) 555-0108 | https://sixthandvine.example | 78701 |
| Mesquite Moon Café | 4.3 | 141 | (512) 555-0110 | https://mesquitemoon.example | 78751 |
| Arbor Lane Café | 4.6 | 59 | (512) 555-0109 | https://arborlane.example | 78703 |
The same rows are in the run's dataset, which Apify can export as JSON, CSV or Excel, and in OUTPUT.csv in the run's key-value store. This is the real output of a local run of the Actor on 4 October 2026.
What it cost
Dataset Transformer charges per row it writes: $1 per 1,000 rows on Apify's Free plan, less on paid plans, plus a fee for each run start. This example wrote 7 rows: 7 × $0.001 = $0.007, plus the run start. The 4 filtered rows and the 4 duplicates were free. The Store page always shows the current price.
Run it after every scrape
Apify can start Dataset Transformer each time your scraper finishes, with no code:
- Open the scraper (or its saved task) in Apify Console and go to its Integrations tab.
- Add Dataset Transformer and set Start when to run succeeded.
- In the input, set Apify dataset to
{{resource.defaultDatasetId}}. Apify replaces it with the dataset of the run that just finished. - Add the filters, duplicate keys, sorting and fields from the input above, and Also save as a file: CSV.
- Save, then test it on one of the scraper's past runs from the integration's test menu.
Keep Apify dataset filled in. If it's empty, the run cleans the Actor's pre-filled example file instead of your scraper's results, and charges for those rows.
To set it up through Apify's API instead, as a webhook on the ACTOR.RUN.SUCCEEDED event, see "Run it automatically after your scraper finishes" in the Actor's Store page. Apify's own guide to chaining Actors is Actor-to-Actor integrations.
Limits
- One source per run: one of your Apify datasets, or one CSV, JSON, JSON Lines or Excel file from a link.
- Up to 1,000,000 rows read per run, and files up to 200 MB.
- It reads only datasets your own Apify account can access.
The other data tools: Dataset Diff returns only the rows that are new, changed or removed since a scraper's last run, and Join Datasets adds fields from a second dataset by a shared key. All of them are on the Apify page.