Literature reviews, citation analysis and reference cleanup all start the same way: you need a table of papers with titles, authors, journals, dates and citation counts. Academic search pages are built for reading one result at a time, not for exporting hundreds of records, and copying metadata by hand is slow and error prone.
Research Paper Search by Hay Equipos takes topics or DOIs and returns one structured row per paper. The data comes from Crossref, the registry that issues DOIs for most of the world's publishers, through its open public API. There is no page scraping involved, so runs are steady and there are no captchas.
What you get back
One row per work. Example with illustrative values:
{
"query": "large language models in education",
"source": "search",
"doi": "10.1234/example.2025.0042",
"doiUrl": "https://doi.org/10.1234/example.2025.0042",
"title": "Using Large Language Models for Feedback in Higher Education",
"type": "journal-article",
"journal": "Journal of Example Education Research",
"publisher": "Example Academic Press",
"publishedDate": "2025-03-14",
"year": 2025,
"authors": ["A. Author", "B. Author"],
"authorCount": 2,
"abstract": "This study examines ...",
"citationCount": 18,
"referenceCount": 54,
"subjects": ["Education"],
"issn": ["1234-5678"],
"openLicense": true,
"funders": [],
"publisherUrl": "https://journal.example.org/articles/42"
}
Rows also include subtitle, firstAuthorAffiliation, isbn, volume, issue, pages, language, licenseUrls, fullTextLinks and relevanceScore. citationCount is the number of other Crossref works that cite the paper. openLicense is true when the work carries a Creative Commons license. The same paper is never saved twice in one run.
Step by step in the Apify Console
- Open the actor on the Apify Store (link at the end) and click Try for free. A free Apify account is enough.
- Fill Search queries with topics, one per line, or DOIs to look up with DOIs or DOI links. You can use both in one run. Queries match titles, authors, journals and other bibliographic text.
- Narrow the results with Published from and Published until (formats
2024,2024-06or2024-06-30) and Work types (journal article, conference paper, book chapter, book, preprint, report, dissertation, dataset and more). - Pick Sort by: relevance, newest first or most cited first.
- Optionally turn on Only works with an abstract or Only works with a license, and switch off Include abstracts for smaller rows.
- Set Results per query (default 100, up to 10,000) and Maximum papers (default 1,000).
- Optionally add Your contact email. Crossref serves requests that include one from a faster pool. It is sent only to Crossref and is not saved in the output.
- Click Start, then export the results as CSV, Excel or JSON.
How the sorting works
"Newest first" and "most cited first" do not sort the whole of Crossref. They take the most relevant matches for your query (five times the results you ask for, at least 200 and at most 1,000) and reorder those. This is deliberate: sorting the entire index by citations ignores relevance and returns papers that are off topic.
Calling it from code
With curl, using the synchronous endpoint that returns the rows directly:
curl -X POST \
"https://api.apify.com/v2/acts/pistachio_implementation~research-paper-search/run-sync-get-dataset-items" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"queries": ["large language models in education"],
"fromDate": "2024",
"types": ["journal-article"],
"sortBy": "most-cited",
"maxResultsPerQuery": 50
}'
With Python and the apify-client package:
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("pistachio_implementation/research-paper-search").call(
run_input={
"queries": ["CRISPR off target effects"],
"dois": ["10.1038/nature14539"],
"fromDate": "2022",
"onlyWithAbstract": True,
"sortBy": "relevance",
"maxResultsPerQuery": 200,
}
)
for paper in client.dataset(run["defaultDatasetId"]).iterate_items():
print(paper["year"], paper["citationCount"], paper["title"], paper["doiUrl"])
summary = client.key_value_store(run["defaultKeyValueStoreId"]).get_record("RUN_SUMMARY")
print(summary["value"]["errors"]) # DOIs not found, queries with no matches
Pricing
Pay per event, with no subscription and no platform usage charge on top:
| Event | Price |
|---|---|
| Paper saved | $0.001 per paper ($1.00 per 1,000 papers) |
There is no start fee. DOIs that are not found and duplicates are free. For example, 200 papers on a topic cost $0.20, and looking up 1,000 DOIs costs $1.00. You can set a maximum charge per run and the actor stops cleanly when it is reached.
Limits and what it does not do
- Abstracts exist only when the publisher deposited one with Crossref. Many older works and some large publishers have none. Use "Only works with an abstract" if you need them.
- Citation counts are Crossref's own citation links. They are usually lower than counts from services that also index web pages and theses.
- Speed: without a contact email, Crossref's public pool allows about one request per second, and each request returns up to 200 papers. With a contact email it is faster.
- Relevance ranking is Crossref's. Results far down a very broad query get looser, so narrow the query or add a date range and type filter.
- Up to 10,000 results per query, which is Crossref's paging limit.
- No full text download. The actor returns links. Access depends on the publisher and the license.
- Only works registered with a DOI are covered. Loose PDFs on the web are not.
- No author profiles. Author names appear only as the paper's published byline.
The actor is an independent tool and is not affiliated with Crossref. Use the data in line with Crossref's terms and each publisher's license.
Try it on the Apify Store: https://apify.com/pistachio_implementation/research-paper-search
Top comments (0)