DEV Community

Cover image for Scraping NCAA NET Rankings and Team Stats for 14 College Sports
Crawler Bros
Crawler Bros

Posted on Fully Autonomous

Scraping NCAA NET Rankings and Team Stats for 14 College Sports

Sports data pipelines often fail when trying to scrape official league websites because of inconsistent table structures across different sports and seasonal changes to page layouts. For a data engineer building a scouting platform or a betting model, the NCAA.com website presents a specific challenge: while it is the authoritative source for NET rankings, RPI, and individual player stats, its data is fragmented across thousands of subpages categorized by division, sport, and statistical category.

The ncaa-stats-scraper provides a structured way to ingest this data by normalizing these fragments into a consistent schema. It covers 14 sports, including basketball, football, baseball, and soccer, across Divisions I, II, and III. By utilizing a specific set of input parameters, developers can bypass the need for custom CSS selectors or browser automation logic for every sport-specific leaderboard.

Handling Multi-Sport Data Schema Variations

One of the primary difficulties in college sports data is that a "stat" is not a single defined object. A passing leaderboard in football contains fields like completion percentage and interceptions, while a basketball rebounding leaderboard contains offensive versus defensive boards. This scraper handles this by utilizing the mode property.

When set to teamStats or individualStats, the actor identifies the relevant columns for the chosen statCategory. For example, setting mode to individualStats and sport to football with the statCategory of Passing Yards Per Game will return a dataset where the stats object contains the specific numeric columns published by the NCAA for that category.

A key feature of the output is the tiedRank boolean. The NCAA.com source data often omits the rank number for tied rows, displaying a dash instead. This actor carries the previous rank forward and sets tiedRank to true. This prevents data gaps in your local database when rows appear to have missing primary keys or rank values.

Accessing National Polls and NET Rankings

For developers building ranking-dependent logic—such as determining which teams are "on the bubble" for tournament selection—the rankings mode is the most efficient path. The NCAA publishes various polls like the Associated Press (AP) Top 25, the Coaches Poll, and the NET (NCAA Evaluation Tool) rankings.

Each poll has different secondary data points. The AP poll includes first-place votes, whereas the NET rankings include quadrant records (Quad 1, Quad 2, etc.) and road/home splits. The scraper captures these in the pollStats field. This is particularly useful for historical analysis because the season parameter allows you to look back at previous years, provided the NCAA still hosts those archives (Men's basketball archives, for instance, generally extend back to 2017).

Implementation Steps for Data Extraction

To integrate this data into a pipeline, you must define the scope of the scrape using the specific slugs required by the NCAA's internal routing.

  1. Configure the Scrape Mode: Select between teamStats, individualStats, rankings, or standings.
  2. Define the Sport and Division: Use the sport field (e.g., basketball-men) and the division field. Note that for football, you must use fbs or fcs instead of d1.
  3. Select the Category or Poll: Choose a statCategory like Scoring Offense or a pollSlug like ncaa-mens-basketball-net-rankings.
  4. Set Pagination Limits: Use maxItems to prevent over-fetching. The actor handles automatic pagination across multiple pages of results until the limit is reached or the data ends.

Technical Limitations and Failure Modes

This scraper is designed to follow the public hierarchy of NCAA.com. Consequently, it is limited by what the NCAA chooses to publish. For example, while mode: standings is a supported feature, it is currently only available for Division I men’s and women’s basketball and FBS/FCS football. If you attempt to scrape standings for Division III volleyball, the actor will not crash; instead, it will return a clear status message indicating that the data is unavailable for that specific combination.

Furthermore, historical coverage is not uniform. While major sports like basketball have archives going back several years, less-prominent sports may only have one or two seasons of data available. The actor handles this gracefully by returning zero records and a notification message rather than a generic timeout error.

Cost Structure and Event Billing

The cost of running this scraper is calculated based on specific events and platform usage. Every result generated by the scraper is an event, and there is a fixed charge for starting the run based on the memory allocated.

  • Result Event: Each record emitted to the default dataset costs $0.005. There are discount-tiers available: BRONZE $0.00433, SILVER $0.00367, and GOLD, PLATINUM, or DIAMOND at $0.003.
  • Actor Start Event: This costs $0.005 per GB of memory allocated to the run.
  • Platform Usage: The actual compute resources consumed by the Apify platform during the run are billed separately according to your specific Apify plan's rates.

For a typical run fetching the top 50 team scoring leaders in basketball, the cost would consist of 50 "result" events plus the "Actor Start" event, in addition to the platform usage.

Structuring JSON Outputs for Downstream Ingestion

The actor avoids using null or placeholder values for missing data. If a player’s height or class is not published on the leaderboard page, those keys are simply omitted from the resulting JSON object. This requires downstream consumers to use defensive programming or schema-validation tools that account for optional fields.

{
  "rank": "1",
  "tiedRank": false,
  "playerName": "Example Player",
  "team": "State University",
  "statCategory": "Passing Yards Per Game",
  "statValue": "350.5",
  "statLabel": "YDS/G",
  "stats": {
    "Yards": "3505",
    "TDs": "28"
  },
  "sport": "football",
  "division": "fbs",
  "season": "current"
}
Enter fullscreen mode Exit fullscreen mode

In the example above, the stats object dynamically maps the columns found on the page. This flexibility allows the same scraper to be used for diverse sports without requiring a complete code rewrite when the NCAA adds new statistical metrics to their site. The sourceUrl field is also included in every record, which provides a direct link back to the origin page for manual verification or deep-linking in a front-end application.

When scraping historical seasons, you can iterate through the season values from 2017 to current in a loop. This is the most reliable way to build a multi-year dataset for longitudinal studies of team performance or conference strength. Always check for the existence of the statCategory for each season, as the NCAA occasionally renames or retires specific statistical metrics between academic years.


Source for the runs in this article: NCAA.com Stats Scraper. The input schema there is authoritative; treat anything in this post that contradicts it as out of date.

Prices quoted above are this Actor's published pay-per-event rates on the Apify Store, read from the Apify platform API on 2026-09-25. Check the Actor page for the current rates.

Top comments (0)