DEV Community

Cover image for Structuring Indian Local Directory Records Across 45 Cities
Crawler Bros
Crawler Bros

Posted on Fully Autonomous

Structuring Indian Local Directory Records Across 45 Cities

Building localized B2B datasets for Indian markets usually involves navigating inconsistent directory markup, unstandardized address strings, and aggressive rate limits on regional portals. AskLaila is one of the oldest business directories in India, holding millions of records across 45+ cities ranging from restaurants and healthcare providers to local trade services.

However, extracting structured records from it introduces classic pipeline challenges: handling pagination, standardizing Schema.org LocalBusiness data, dealing with multiple contact numbers, and handling transient 503 errors on datacenter IP ranges.

The AskLaila India Business Directory Scraper handles the extraction, data normalization, and failover routing directly.

Handling AskLaila Data Structures and Cloud Blocks

AskLaila serves server-side rendered pages with embedded Schema.org microdata, but querying it at scale runs into two main issues: regional routing failures and deeply nested field variants across service categories.

The scraper maps the underlying Schema.org microdata (LocalBusiness, Restaurant, Hospital, etc.) into a consistent record structure. It emits:

  • Entity Details: name, category, businessType, breadcrumbs, listingId, and canonicalUrl.
  • Geographic Data: streetAddress, locality, city, postalCode, fullAddress, landmark, latitude, and longitude.
  • Contact Points: phoneNumbers extracted as clean arrays of digits-only strings.
  • Reputation Metrics: ratingValue, ratingCount, reviewCount, and recommendPercent.
  • Category Attributes: An attributes object mapping service-specific metadata (e.g., AC, Home Delivery, Working Hours, Valet Parking, Credit Cards Accepted).

When running automated pipelines on datacenter hardware, directory pages can occasionally time out or return empty responses. The actor handles this through a multi-tier fallback:

  1. It attempts a direct request.
  2. If blocked or empty, and autoEscalateOnBlock is enabled, it automatically retries through standard Apify Proxy AUTO.
  3. If the first-party endpoint remains unreachable, it can leverage indexedSearch to surface verified Google-indexed AskLaila URLs without failing the pipeline run.

Limitations to Note

This actor cannot perform full-text business-name searches across the entire directory via keywords. If you supply a specific brand name like "Apollo Hospital" or "Domino's" to keywordSearch, AskLaila's search endpoint defaults to fallback city pages instead of matching the specific business. For individual entity lookups, you must supply exact URLs using listingUrls or searchUrls.

Choosing the Right Scrape Mode

The actor accepts five crawl configurations depending on the upstream data you already hold:

Mode Purpose Core Input Parameters
byCategory Query standard category listings within a specific city city, category, optional locality
keywordSearch Query specific service slugs (e.g., pizza restaurants, ac repair) city, keyword
listingUrls Direct detail extraction and review scraping for known profile URLs listingUrls
searchUrls Direct execution over pre-filtered search result URLs searchUrls
indexedSearch Fallback search querying indexed records when first-party routing fails city, category or keyword

Configuring a Run: Step-by-Step

Here is a standard extraction workflow targeting dining establishments in Bangalore's Koramangala area, enriched with review content and filtered by minimum ratings.

1. Define the Run Configuration

Set up the JSON payload using byCategory mode. By setting fetchListingDetails to true, the actor will make a secondary request to each business's profile page to pull deep attributes, coordinates, and visible reviews.

{
  "mode": "byCategory",
  "city": "Bangalore",
  "category": "restaurants",
  "locality": "Koramangala",
  "fetchListingDetails": true,
  "includeReviews": true,
  "minRating": 4,
  "minReviewCount": 5,
  "maxItems": 100,
  "autoEscalateOnBlock": true
}
Enter fullscreen mode Exit fullscreen mode

2. Execute via the Apify Python Client

You can run the actor programmatically within an orchestration pipeline using Python:

from apify_client import ApifyClient

client = ApifyClient("YOUR_API_TOKEN")

run_input = {
    "mode": "byCategory",
    "city": "Bangalore",
    "category": "restaurants",
    "locality": "Koramangala",
    "fetchListingDetails": True,
    "includeReviews": True,
    "minRating": 4,
    "maxItems": 100,
    "autoEscalateOnBlock": True
}

run = client.actor("crawlerbros/asklaila-india-business-directory-scraper").call(run_input=run_input)
dataset_items = client.dataset(run["defaultDatasetId"]).list_items().items

for item in dataset_items:
    print(f"Name: {item.get('name')}")
    print(f"Address: {item.get('fullAddress')}")
    print(f"Phones: {item.get('phoneNumbers')}")
    print(f"Coordinates: {item.get('latitude')}, {item.get('longitude')}")
    print("---")
Enter fullscreen mode Exit fullscreen mode

3. Normalize the Output Schema

The output dataset omits null fields, returning populated keys for each listing. Here is an example of an emitted record:

{
  "name": "Hallimane",
  "category": "Restaurant",
  "businessType": "Restaurant",
  "breadcrumbs": "Bangalore > Restaurants > Malleswaram > Hallimane",
  "phoneNumbers": [
    "08023348888",
    "08023348889"
  ],
  "streetAddress": "3rd Cross, Sampige Road",
  "locality": "Malleswaram",
  "city": "Bangalore",
  "postalCode": "560003",
  "fullAddress": "3rd Cross, Sampige Road, Malleswaram, Bangalore - 560003",
  "landmark": "Near Circle Maramma Temple",
  "latitude": 13.0031,
  "longitude": 77.5703,
  "ratingValue": 4.2,
  "ratingCount": 128,
  "reviewCount": 45,
  "recommendPercent": 85,
  "cuisine": "South Indian, North Indian",
  "paymentAccepted": "Cash, Credit Card, Debit Card",
  "attributes": {
    "Air Conditioned": "Yes",
    "Home Delivery": "Yes",
    "Vegetarian": "Pure Veg"
  },
  "reviews": [
    {
      "authorName": "Ramesh K.",
      "ratingValue": 5,
      "datePublished": "2023-11-12",
      "reviewBody": "Authentic traditional food with quick service."
    }
  ],
  "listingUrl": "https://www.asklaila.com/listing/Bangalore/malleswaram/hallimane/1OjwP4uM/",
  "listingId": "1OjwP4uM",
  "dataSource": "asklaila",
  "sourcePageType": "detail"
}
Enter fullscreen mode Exit fullscreen mode

Understanding Event-Based Run Costs

The actor operates on a pay-per-event pricing model. There are two primary billing events:

  1. Actor Start (apify-actor-start): Billed at $0.005 per GB of memory allocated to the run when the Actor starts.
  2. Result (apify-default-dataset-item): Billed per dataset item emitted. The base rate on the FREE tier is $0.005 per event. Depending on your account's discount tier, this unit cost decreases:
    • BRONZE: $0.00433 per event
    • SILVER: $0.00367 per event
    • GOLD / PLATINUM / DIAMOND: $0.003 per event

In addition to event charges, runs consume standard Apify platform usage billed separately at your account's plan rates. If you need to enrich large volumes of data while controlling dataset emissions, configuring minRating and minReviewCount upstream prevents low-quality leads from generating billable dataset items.

A useful next step for downstream pipelines is implementing deduplication on the listingId field, particularly when running broad scrapes across overlapping localities in large metro areas.


Runs in this article used AskLaila India Business Directory Scraper. Its README is the reference for input fields and output structure; this post is only one path through them.

Prices quoted above are this Actor's published pay-per-event rates on the Apify Store, read from the Apify platform API on 2026-09-25. Check the Actor page for the current rates.

Top comments (0)