White Hat Harbor

Clean Apify scraper results into a CSV

A scraper's raw results usually need the same clean-up before they go into a spreadsheet or CRM: drop the rows you don't want, remove repeats, pick and rename the columns. This page walks through one example with my Apify Actor, Dataset Transformer, from the raw rows to the CSV.

The scraper's results

A Google Maps-style scraper ran two searches over the same area, "coffee shop in Austin, TX" and "cafe in Austin, TX". Places that match both searches come back twice, and each row has 24 fields, most of them nested or not needed in a spreadsheet: opening hours, image lists, the search rank and the scrape time. The places, phone numbers (555-01xx) and websites (.example) are made up for this example.

The goal: one CSV with one row per open place rated 4.3 or higher that has a phone number, the most reviewed first, with plain column names.

Before: 15 rows

Six of the 24 fields, and what happens to each row:

#titleWhat happenstotalScorereviewsCountphonepermanentlyClosedsearchString
1Juniper & Oak CoffeeKept4.71284(512) 555-0101falsecoffee shop in Austin, TX
2Lantern Street RoastersKept4.6902(512) 555-0102falsecoffee shop in Austin, TX
3Copper Kettle CaféKept (no website: empty cell)4.4377(512) 555-0103falsecoffee shop in Austin, TX
4Night Owl Espresso BarFiltered out: rated below 4.34.1215(512) 555-0104falsecoffee shop in Austin, TX
5Riverbend Bakery & CoffeeKept4.82140(512) 555-0105falsecoffee shop in Austin, TX
6The Grind HouseFiltered out: permanently closed4.5640(512) 555-0106truecoffee shop in Austin, TX
7Bluebonnet Brew Co.Filtered out: no phone4.588nullfalsecoffee shop in Austin, TX
8Sixth & Vine CoffeeKept4.5312(512) 555-0108falsecoffee shop in Austin, TX
9Arbor Lane CaféKept4.659(512) 555-0109falsecoffee shop in Austin, TX
10Riverbend Bakery & CoffeeDuplicate of 5 (same placeId)4.82140(512) 555-0105falsecafe in Austin, TX
11Mesquite Moon CaféKept4.3141(512) 555-0110falsecafe in Austin, TX
12Juniper & Oak CoffeeDuplicate of 14.71284(512) 555-0101falsecafe in Austin, TX
13Pecan Tree Tea & CoffeeFiltered out: rated below 4.34.2403(512) 555-0111falsecafe in Austin, TX
14Lantern Street RoastersDuplicate of 24.6902(512) 555-0102falsecafe in Austin, TX
15Copper Kettle CaféDuplicate of 34.4377(512) 555-0103falsecafe in Austin, TX

The input

Three filters, duplicates removed by the place's ID, sorted by review count, and twelve columns kept, eight of them renamed (old -> new). Paste it into the Actor's JSON input, or fill in the same values in the form, with your scraper run's dataset in Apify dataset:

{
  "datasetId": "YOUR-SCRAPER-DATASET-ID",
  "filters": ["permanentlyClosed != true", "totalScore >= 4.3", "phone is not empty"],
  "dedupeBy": ["placeId"],
  "sortBy": ["reviewsCount desc"],
  "fields": ["title -> name", "categoryName -> category", "totalScore -> rating", "reviewsCount -> reviews",
             "phone", "website", "street", "city", "postalCode -> zip",
             "location.lat -> lat", "location.lng -> lng", "url -> googleMapsUrl"],
  "exportFormat": "csv"
}

Filters, duplicates and sorting use the scraper's own field names (totalScore, placeId). The new names only appear in the output. location.lat reaches into the nested location object.

After: 7 rows

The run's status line: "7 rows written from 15 rows read; 4 filtered out; 4 duplicates removed; file saved as OUTPUT.csv". Six of the twelve columns (the others are category, street, city, lat, lng and googleMapsUrl):

nameratingreviewsphonewebsitezip
Riverbend Bakery & Coffee4.82140(512) 555-0105https://riverbendbakery.example78746
Juniper & Oak Coffee4.71284(512) 555-0101https://juniperoak.example78702
Lantern Street Roasters4.6902(512) 555-0102https://lanternstreet.example/menu?utm_source=gmb78704
Copper Kettle Café4.4377(512) 555-010378741
Sixth & Vine Coffee4.5312(512) 555-0108https://sixthandvine.example78701
Mesquite Moon Café4.3141(512) 555-0110https://mesquitemoon.example78751
Arbor Lane Café4.659(512) 555-0109https://arborlane.example78703

The same rows are in the run's dataset, which Apify can export as JSON, CSV or Excel, and in OUTPUT.csv in the run's key-value store. This is the real output of a local run of the Actor on 4 October 2026.

What it cost

Dataset Transformer charges per row it writes: $1 per 1,000 rows on Apify's Free plan, less on paid plans, plus a fee for each run start. This example wrote 7 rows: 7 × $0.001 = $0.007, plus the run start. The 4 filtered rows and the 4 duplicates were free. The Store page always shows the current price.

Run it after every scrape

Apify can start Dataset Transformer each time your scraper finishes, with no code:

  1. Open the scraper (or its saved task) in Apify Console and go to its Integrations tab.
  2. Add Dataset Transformer and set Start when to run succeeded.
  3. In the input, set Apify dataset to {{resource.defaultDatasetId}}. Apify replaces it with the dataset of the run that just finished.
  4. Add the filters, duplicate keys, sorting and fields from the input above, and Also save as a file: CSV.
  5. Save, then test it on one of the scraper's past runs from the integration's test menu.

Keep Apify dataset filled in. If it's empty, the run cleans the Actor's pre-filled example file instead of your scraper's results, and charges for those rows.

To set it up through Apify's API instead, as a webhook on the ACTOR.RUN.SUCCEEDED event, see "Run it automatically after your scraper finishes" in the Actor's Store page. Apify's own guide to chaining Actors is Actor-to-Actor integrations.

Limits

The other data tools: Dataset Diff returns only the rows that are new, changed or removed since a scraper's last run, and Join Datasets adds fields from a second dataset by a shared key. All of them are on the Apify page.