The Architectural Challenge of Ephemeral Snapchat CDN Links
Extracting data from Snapchat Spotlight is an effective way to analyze trends, gather engagement metrics, and archive social media content. However, developers integrating these datasets into their backend pipelines immediately run into a major hurdle: the direct CDN links provided in the videoUrl field are temporary. Snapchat's media servers secure their content using signed URLs that include expiration tokens. If you store the raw video link in a database and attempt to fetch it hours or days later, the request will fail with an HTTP 403 Forbidden or 410 Gone error.
To build a resilient ingestion pipeline, you must download the video binaries immediately after the scraper completes its execution. By leveraging the snapchat-spotlight-video-downloader on the Apify platform, you can retrieve direct CDN URLs, full engagement stats, transcripts, and AI-generated metadata. However, because these URLs are short-lived, your downloader architecture must be highly optimized.
This article details how to manage the lifecycle of these ephemeral URLs, how to bypass standard platform timeouts when scraping large lists of links, how to structure your inputs and outputs, and how to write high-throughput code that handles download queues before tokens expire.
How to Get Direct Snapchat Spotlight Video URLs?
You can get direct Snapchat Spotlight CDN video URLs by submitting a list of target links to the Snapchat Spotlight Video Downloader. This Actor processes public Spotlight pages to extract direct video addresses, creator profiles, engagement metrics, and transcripts without requiring a login or cookies. Standard profile links, specific Spotlight post links, and short links are all supported as input.
The input schema requires the spotlightUrls array. This array accepts standard Spotlight links, creator-specific URLs, and mobile short links. The Actor automatically resolves redirecting short links like https://t.snapchat.com/{id} to their canonical location before scraping.
The following JSON example shows a structured input that targets multiple Spotlight posts. It configures the run to extract the maximum amount of contextual information, including AI-generated descriptions and WebVTT audio transcripts:
{
"spotlightUrls": [
"https://www.snapchat.com/spotlight/W7_EDlXWTBiXAEEniNoMPwAAYZmRwYmVqcm5wAZ8Kxg_vAZ8Kxgn9AAAAAQ",
"https://www.snapchat.com/spotlight/W7_EDlXWTBiXAEEniNoMPwAAYZG5iaHJuY2phAZ8KF10GAZ8KETc1AAAAAQ"
],
"includeTranscript": true,
"includeAiMetadata": true,
"includeComments": false,
"includeRelatedStories": false
}
By explicitly setting these parameters, you control the size and detail of the output records. The proxyConfiguration field is optional because public Spotlight pages are accessible without proxies in most geographical regions.
What Output Fields are Available for Spotlight Videos?
The output fields available for Spotlight videos include the direct CDN video URL, high-resolution thumbnail images, video height and width dimensions, full engagement metrics, creator usernames, and structured metadata. If enabled, the schema also supplies LLM-generated titles and transcripts. These fields allow your downstream systems to archive both the raw media files and their context.
The Actor outputs a structured JSON record for each successfully scraped Spotlight URL, containing core video attributes and any requested metadata arrays. This structure includes the unique snapId, the original inputUrl, the direct videoUrl, creator profile details, and precise engagement metrics.
The following output shows the exact structure of a successfully processed record, matching the fields and structure returned by the platform:
{
"snapId": "W7_EDlXWTBiXAEEniNoMPwAAYZG5iaHJuY2phAZ8KF10GAZ8KETc1AAAAAQ",
"inputUrl": "https://www.snapchat.com/spotlight/W7_EDlXWTBiXAEEniNoMPwAAYZG5iaHJuY2phAZ8KF10GAZ8KETc1AAAAAQ",
"videoUrl": "https://bolt-gcdn.sc-cdn.net/v/udlVbjK0wdjHbPx1FRph0.27.IRZXSOY",
"thumbnailUrl": "https://bolt-gcdn.sc-cdn.net/v/udlVbjK0wdjHbPx1FRph0.256.IRZXSOY",
"caption": "Another Spotlight Snap brought to you by Snapchat",
"creatorUsername": "waqarzaka",
"profileUrl": "https://www.snapchat.com/@waqarzaka",
"creatorDisplayName": "Waqar Zaka",
"durationSeconds": 130.23,
"width": 540,
"height": 960,
"shareCount": 0,
"likeCount": 3,
"commentCount": 0,
"recommendCount": 0,
"uploadedAt": "2026-06-27T17:12:08.245000+00:00",
"scrapedAt": "2026-06-28T06:43:34.588497+00:00",
"isAttributed": true,
"textMetadataKeywords": ["bitcoin", "hacker", "crypto", "cryptocurrency", "blockchain"],
"contextCards": [
{
"contextType": 3,
"title": "Waqar Zaka",
"subtitle": "waqarzaka",
"url": "https://www.snapchat.com/@waqarzaka",
"thumbnailType": 0,
"hasBadge": true,
"id": "36d0088b-39d5-4765-a237-bb147a292272"
}
],
"detectedLanguages": [
{ "languageCode": "hi", "score": 0.910262405872345 }
]
}
The presence of the videoUrl allows your downstream tasks to retrieve the binary payload. Notice how isAttributed identifies if the snap belongs to the main platform feed.
How to Page Through Results for Many Spotlight Videos?
You can page through results for many Spotlight videos by using the offset and limit query parameters against the default dataset API endpoint. This approach allows you to fetch records in sequential, manageable batches rather than pulling the entire dataset into memory at once. It protects your local system from memory exhaustion when handling datasets containing hundreds of scraped videos.
The core endpoint for fetching items is https://api.apify.com/v2/datasets/<DATASET_ID>/items. When querying this endpoint, passing the query parameter clean=true is recommended to omit platform-specific metadata columns like crawler run identifiers.
The following Python code uses an asynchronous request pattern to paginate through a dataset and extract all returned records for downstream storage:
import asyncio
import aiohttp
async def fetch_dataset_page(session, dataset_id, offset, limit):
url = f"https://api.apify.com/v2/datasets/{dataset_id}/items"
params = {
"offset": offset,
"limit": limit,
"clean": "true"
}
async with session.get(url, params=params) as response:
if response.status == 200:
return await response.json()
return []
async def fetch_all_results(dataset_id, limit=100):
async with aiohttp.ClientSession() as session:
offset = 0
all_items = []
while True:
items = await fetch_dataset_page(session, dataset_id, offset, limit)
if not items:
break
all_items.extend(items)
offset += limit
return all_items
This asynchronous approach reduces overhead and quickly aggregates results. It avoids the synchronous blockers common in large-scale data retrieval tasks.
What if a Run Exceeds the 300 Second Synchronous Cap?
If an Actor run takes longer than the 300-second platform cap, you must initiate the run asynchronously and poll the run's status or use webhooks. The synchronous run endpoint hard-caps at 300 seconds and returns an HTTP 408 timeout past that threshold. Initiating the run asynchronously allows your application to handle jobs containing hundreds of target links without timing out.
When you issue a POST request to /v2/acts/crawlerbros~snapchat-spotlight-video-downloader/runs, the platform immediately returns a JSON response containing the run's metadata and status. This allows your calling application to release the connection and poll the status endpoint or wait for a webhook payload.
The following Python example shows how to launch an asynchronous run and poll its status until it completes, ensuring your integration handles heavy jobs without encountering network timeouts:
import asyncio
import aiohttp
import sys
async def monitor_run(api_token, actor_id, run_input):
headers = {"Authorization": f"Bearer {api_token}"}
start_url = f"https://api.apify.com/v2/acts/{actor_id}/runs"
async with aiohttp.ClientSession(headers=headers) as session:
async with session.post(start_url, json=run_input) as response:
if response.status != 201:
print("Failed to start the Actor run asynchronously.")
sys.exit(1)
run_data = await response.json()
run_id = run_data["data"]["id"]
print(f"Asynchronous run started. Run ID: {run_id}")
status_url = f"https://api.apify.com/v2/actor-runs/{run_id}"
while True:
async with session.get(status_url) as response:
if response.status == 200:
status_data = await response.json()
status = status_data["data"]["status"]
print(f"Current Status: {status}")
if status in ["SUCCEEDED", "FAILED", "ABORTED", "TIMED_OUT"]:
return status_data["data"]
await asyncio.sleep(20)
By checking the run's progress, your pipeline remains responsive and can process lists of any size. Note that inputs defined via the API must always contain an explicit JSON object, as schema-level prefill parameters only apply in the Console UI and are ignored on API calls.
Predicting Token Expiration From Snapchat CDN URLs
Because Snapchat CDN URLs are signed with time-based tokens, downloading them requires a clear understanding of when they will stop working. Instead of blindly trying to download links and catching errors, you can reverse-engineer the expiration timestamp directly from the URL query parameters. This allows your script to prioritize links that are closest to expiring.
Typically, Snapchat's CDN servers (like bolt-gcdn.sc-cdn.net) append query parameters containing security tokens and expiration timestamps. In many CDN architectures, parameters such as expiry or expires contain an epoch timestamp representing the exact second the link becomes invalid.
The following Python code parses a signed Snapchat CDN URL, extracts the expiration timestamp, and calculates the remaining lifespan of the token in seconds:
from urllib.parse import urlparse, parse_qs
import time
def estimate_link_lifespan(cdn_url: str) -> float:
parsed = urlparse(cdn_url)
query_params = parse_qs(parsed.query)
# Check common CDN expiration parameter names
expiry_keys = ["expiry", "expires", "exp", "et"]
for key in expiry_keys:
if key in query_params:
try:
expiry_timestamp = int(query_params[key][0])
# If timestamp is in milliseconds, convert to seconds
if expiry_timestamp > 9999999999:
expiry_timestamp /= 1000.0
remaining_time = expiry_timestamp - time.time()
return max(0.0, remaining_time)
except ValueError:
continue
# If no explicit parameter is matched, return a default safe margin
return 3600.0
By extracting this data, you can build a priority queue where records are sorted by their remaining lifetime, ensuring that no link expires before your downloader can fetch its payload.
High-Throughput Async Video Downloads Before Token Expiration
When handling hundreds of Spotlight records, downloading them sequentially using standard libraries will lead to expired links. The solution is an asynchronous, high-throughput downloader that limits concurrent connections to avoid rate limiting while maximizing bandwidth usage.
Using asyncio and aiohttp, you can process multiple video streams simultaneously. This ensures that the entire batch of videos is downloaded well within the typical validity window of the CDN tokens.
The following Python script reads the scraped dataset, filters out records lacking a valid videoUrl, and downloads the binaries to local storage:
import asyncio
import aiohttp
from pathlib import Path
async def download_video_stream(session, semaphore, record, output_dir):
video_url = record.get("videoUrl")
snap_id = record.get("snapId")
creator = record.get("creatorUsername")
if not video_url or not snap_id:
return
output_path = Path(output_dir) / f"{creator or 'unknown'}_{snap_id}.mp4"
async with semaphore:
try:
async with session.get(video_url, timeout=60) as response:
if response.status == 200:
with open(output_path, "wb") as f:
while True:
chunk = await response.content.read(65536)
if not chunk:
break
f.write(chunk)
print(f"Downloaded: {output_path.name}")
else:
print(f"HTTP Error {response.status} for snap {snap_id}")
except Exception as e:
print(f"Failed to download snap {snap_id}: {str(e)}")
async def batch_download_videos(records, output_dir, concurrency_limit=10):
Path(output_dir).mkdir(exist_ok=True)
semaphore = asyncio.Semaphore(concurrency_limit)
async with aiohttp.ClientSession() as session:
tasks = [
download_video_stream(session, semaphore, r, output_dir)
for r in records
]
await asyncio.gather(*tasks)
This script controls concurrency via asyncio.Semaphore to protect your local system from network exhaustion, while ensuring that the downloads are finished before the signed URLs expire.
Identifying and Handling Missing View Counts
A common point of failure in social media data pipelines is assuming that all analytical counters are always populated. For Snapchat Spotlight videos, the viewCount field exhibits a distinct behavior. For very recent uploads, Snapchat's public platform returns a sentinel value of -1 because the analytics systems have not yet finalized the view count.
To avoid displaying confusing negative values or crashing downstream visualization databases, the Actor omits the viewCount field entirely from the output object when this sentinel value is detected. Your backend data ingestion system must check for this key's presence and apply a default value or an "Omitted" string.
The following Python example shows how to parse incoming records safely, identifying and handling records where the viewCount is missing:
def process_analytical_metrics(record):
snap_id = record.get("snapId", "N/A")
creator = record.get("creatorUsername", "unknown")
# Safely extract viewCount, providing 'Omitted' if the key is absent
view_count = record.get("viewCount", "Omitted")
# Extract engagement metrics
shares = record.get("shareCount", 0)
likes = record.get("likeCount", 0)
comments = record.get("commentCount", 0)
print(f"Processing Snap: {snap_id} by @{creator}")
print(f"Views: {view_count} | Likes: {likes} | Shares: {shares} | Comments: {comments}")
return {
"snap_id": snap_id,
"creator": creator,
"views": view_count,
"likes": likes,
"shares": shares,
"comments": comments
}
This handling strategy prevents database type errors when loading data into columns that expect integer values, while preserving the accuracy of your reporting tools.
Snapchat Spotlight Downloader Limitations and Platform Caveats
Designing a resilient pipeline around the snapchat-spotlight-video-downloader requires understanding its technical limits.
First, the direct CDN links in the videoUrl field contain active signatures. They will expire quickly. You must download the files immediately rather than expecting the URLs to persist in long-term storage.
Second, the structural metadata returned by the scraper is highly dependent on what Snapchat exposes. Some Spotlight posts do not contain video assets (e.g., static images or interactive filters), meaning that the width and height dimensions will be completely missing from the output object. Transcripts are also subject to this limitation; if a video lacks audio metadata or has no track available, the transcript array will be empty.
Third, you must account for the platform's data retention limitations. Unnamed storages on the Apify platform expire. If you are on the free plan, only your 10 most recent runs are retained, and they are fully deleted after 4 months. To prevent data loss, you must save your results to named storages, which are permanently exempt from this cleanup logic.
Furthermore, request queues have operational limits. A request queue can only be processed by one Actor or task run at a time. Trying to run parallel container tasks off a single, shared request queue to speed up extraction will not work. Finally, Apify provides no native AWS S3 or Slack integrations. Moving downloaded video binaries to an S3 bucket or posting automated summaries to a Slack channel requires you to write custom integration code inside your own backend infrastructure or route webhooks through services like Make, Zapier, or n8n.
Calculating Run Costs Under the Pay-Per-Event Model
The cost of running this Actor is calculated using a pay-per-event pricing model. Under this model, you are charged an event-based fee when specific actions are triggered during a run, in addition to standard platform usage. The event-based portion of the cost scales with the number of events your run emits, while the platform-usage portion scales with the system resources consumed. On top of these, you also pay the Apify platform usage the run consumes, which is billed separately at your Apify plan's rates.
The event charges are structured as follows:
- "Actor Start" (apify-actor-start): Billed once per run at a rate of $0.005 per GB of memory allocated to the run.
- "result" (apify-default-dataset-item): Billed at a base rate of $0.002 per event for each single result record saved to your default dataset.
Depending on your Apify user account, different discount-tier prices apply for the "result" event. The exact rates for these tiers are:
- FREE: $0.002 per event
- BRONZE: $0.00167 per event
- SILVER: $0.00133 per event
- GOLD: $0.001 per event
- PLATINUM: $0.001 per event
- DIAMOND: $0.001 per event
Your final cost is a combination of these fixed event fees and the variable platform usage consumed by your crawler run. Setting input flags like includeComments or includeRelatedStories increases the size of your dataset entries, which will influence the resources used during the run.
Checked against the Actor's input schema and Apify docs on 2026-09-26.
The Actor's README is the source of truth for its inputs, outputs and limits. Need a hand wiring this into your stack? Email info@crawlerbros.com
Top comments (0)