Tracking hiring trends in South Africa requires pulling live structured data from localized job boards. CareerJunction host tens of thousands of active postings across 26 job categories, but extracting these listings programmatically presents two technical roadblocks: strict HTTP client blocking that 302-redirects non-browser requests to /not-found, and short-lived Amazon S3 image URLs that break logo assets within 24 hours of scraping.
The CareerJunction Jobs Scraper extracts live records from this platform while handling browser-fingerprint challenges and image rehosting automatically.
Bypassing the /not-found Redirect on CareerJunction Listings
Standard HTTP requests sent via basic curl or lightweight Python libraries like requests fail when hitting CareerJunction listing URLs. Even if you pass a modern browser User-Agent header, the server's bot-detection infrastructure checks for specific TLS/browser fingerprints. If the client fails this check, the server issues a 302 redirect to a /not-found endpoint, hiding the underlying job data.
To observe this issue, curl a direct job listing URL from a terminal:
curl -I -H "User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64)" \
"https://www.careerjunction.co.za/digital-coordinator-job-2642426.aspx"
Instead of a 200 OK response with HTML content, the server returns:
HTTP/1.1 302 Found
Location: https://www.careerjunction.co.za/not-found
The scraper bypasses this security layer by using browser-like fetch signatures without requiring paid proxies or manual cookie management. This allows direct ingestion of the listing's target content and underlying JSON payloads.
Resolving 24-Hour Expiry on Company Logos
CareerJunction serves employer logos via presigned AWS S3 URLs. When a scraper extracts the raw HTML src attribute, it captures a temporary URL structured like this:
https://cj-employer-logos.s3.amazonaws.com/logo_123.png?AWSAccessKeyId=AKIA...&Expires=1710000000&Signature=...
Because these presigned links expire after roughly 24 hours, storing the raw URL in a database results in broken asset links the next day.
To solve this, the scraper fetches the logo binary during execution, rehosts it in permanent platform storage, and yields two distinct fields in the output dataset:
-
companyLogoUrl: A permanent, rehosted image URL safe for long-term database storage. -
companyLogoUrlOriginal: The raw, presigned S3 link captured directly from CareerJunction.
Structured Field Extraction and Output Schema
CareerJunction packs metadata—including employment type, seniority level, and Employment Equity status—into a single composite string on the listing page (such as "Permanent Senior EE position"). The scraper parses these raw strings into structured fields.
When a field is missing on the source page (e.g., when an employer does not disclose salary information), the key is completely omitted from the output JSON record rather than returning empty strings or null values.
A fully parsed item record contains the following keys:
{
"jobId": "2642426",
"title": "Senior Backend Developer",
"companyName": "Tech Corp",
"companyUrl": "https://www.careerjunction.co.za/companies/tech-corp",
"companyLogoUrl": "https://api.apify.com/v2/key-value-stores/.../records/logo.png",
"companyLogoUrlOriginal": "https://cj-employer-logos.s3.amazonaws.com/...",
"location": "Cape Town",
"country": "South Africa",
"countryCode": "ZA",
"industry": "Information Technology",
"positionText": "Permanent Senior EE position",
"employmentType": "Permanent",
"seniorityLevel": "Senior",
"isEmploymentEquity": true,
"salaryText": "R80 000 - R100 000 per month",
"salaryMin": 80000,
"salaryMax": 100000,
"salaryCurrency": "ZAR",
"salaryPeriod": "Monthly",
"descriptionHtml": "<p>Full job description...</p>",
"descriptionText": "Full job description...",
"datePosted": "2026-03-20T08:00:00.000Z",
"validThrough": "2026-04-20T08:00:00.000Z",
"sourceUrl": "https://www.careerjunction.co.za/senior-backend-developer-job-2642491.aspx",
"recordType": "job",
"scrapedAt": "2026-03-30T10:15:00.000Z"
}
Configuring Search Modes and Filters
The scraper operates in two distinct operational modes via the mode parameter:
-
search: Discovers listings dynamically via keyword queries, server-side region codes, and job categories. -
byUrls: Directly processes an explicit array of CareerJunction listing URLs provided in thejobUrlsinput parameter.
Filtering Options
-
Server-Side Regional Filtering: The
regionparameter accepts specific region keys likewesternCape,gauteng, orworkFromHome. These correspond to CareerJunction's native query parameters for fast server-side target selection. -
Client-Side Substring Matching: The
locationparameter filters strings post-fetch (e.g., matching"Sandton"or"Rosebank"within a broader region). -
Client-Side Attribute Filtering:
employmentTypes,seniorityLevels,employmentEquityOnly,minSalary, andpostedWithinDaysnarrow down results post-extraction. If a listing's rawpositionTextdoes not follow standard parsing conventions, the scraper passes the record through rather than silently dropping data.
Example Input Configuration
This JSON input targets recent IT jobs in the Western Cape with detailed descriptions enabled:
{
"mode": "search",
"keywords": "developer",
"category": "informationTechnology",
"region": "westernCape",
"sortBy": "newest",
"employmentEquityOnly": false,
"minSalary": 50000,
"postedWithinDays": 7,
"fetchFullDetails": true,
"maxItems": 100
}
Running the Scraper Step-by-Step
Step 1: Define Input Schema
Create a JSON configuration file input.json containing your target parameters and search restrictions:
{
"mode": "search",
"keywords": "accountant",
"category": "finance",
"region": "gauteng",
"maxItems": 20
}
Step 2: Execute the Actor
Run the scraper using the Apify JavaScript SDK or Python API client:
from apify_client import ApifyClient
client = ApifyClient("YOUR_API_TOKEN")
run = client.actor("crawlerbros/careerjunction-scraper").call(
run_input={
"mode": "search",
"keywords": "accountant",
"category": "finance",
"region": "gauteng",
"maxItems": 20,
}
)
dataset_items = client.dataset(run["defaultDatasetId"]).list_items().items
for item in dataset_items:
print(f"{item.get('title')} - {item.get('companyName')}")
Step 3: Consume Dataset Output
Retrieve extracted records from the resulting default dataset. Disable fetchFullDetails ("fetchFullDetails": false) if you only require search-level card data (title, company, location, salary text) to increase execution speed.
Event Cost Structure
The scraper operates under a Pay-Per-Event billing model. Charges are calculated based on explicit runtime events:
- Result event (
apify-default-dataset-item): Each result costs $0.005 on the FREE tier ($0.00433 on BRONZE, $0.00367 on SILVER, and $0.003 on GOLD, PLATINUM, and DIAMOND tiers). - Actor Start event (
apify-actor-start): Charged at $0.005 per GB of memory allocated to the run.
Platform usage for the run is billed separately at your Apify plan's rates.
Limitations and Operational Bounds
While the scraper handles job search pages, direct job URLs, and listing details, it cannot scrape dedicated employer/company profile pages under the /companies/... route. CareerJunction protects company profile pages behind an additional bot-challenge layer that blocks direct HTTP collection methods. Accessing company profile data requires direct navigation through standard job detail listings instead.
Runs in this article used CareerJunction Jobs Scraper. Its README is the reference for input fields and output structure; this post is only one path through them.
Prices quoted above are this Actor's published pay-per-event rates on the Apify Store, read from the Apify platform API on 2026-10-06. Check the Actor page for the current rates.
Top comments (0)