DEV Community

Praveen Pasco
Praveen Pasco

Posted on

Tracking vendor release notes without a browser: Notion's loadPageChunk and Zendesk's edited_at

I keep a small text file of dates: the days a few AI writing detectors swapped the model that does the scoring. The reason is boring. If someone runs the same essay through the same detector two weeks apart and gets two different numbers, the first thing I want to know is whether the model changed in between. If it did, the two numbers aren't really comparable.

Keeping that file up to date by hand meant opening a browser on two kinds of pages I can't fetch with curl. This week I finally scripted both, and each one had a small gotcha I didn't expect. Writing them down here mostly so I stop relearning them.

Problem 1: a Notion page that is all JavaScript

GPTZero publishes its model release notes as a public Notion page:

https://gptzero.notion.site/GPTZero-Release-Notes-Model-and-API-6f58686f6381498baef35212463b7da6

Fetch it with curl and you get about 20 KB of HTML whose only visible text is "Notion JavaScript must be enabled in order to use Notion. Please enable JavaScript to continue." The content is loaded afterwards by the Notion client, from an endpoint called loadPageChunk. That endpoint answers plain POSTs for public pages.

The page ID is the 32 hex characters at the end of the URL, reformatted as a dashed UUID:

raw = "6f58686f6381498baef35212463b7da6"
page_id = f"{raw[:8]}-{raw[8:12]}-{raw[12:16]}-{raw[16:20]}-{raw[20:]}"
print(page_id)  # 6f58686f-6381-498b-aef3-5212463b7da6
Enter fullscreen mode Exit fullscreen mode

Then:

curl -s -X POST https://gptzero.notion.site/api/v3/loadPageChunk \
  -H 'Content-Type: application/json' \
  -d '{"pageId":"6f58686f-6381-498b-aef3-5212463b7da6","limit":100,"cursor":{"stack":[]},"chunkNumber":0,"verticalColumns":false}' \
  -o chunk.json -w "%{http_code} %{size_download}\n"
Enter fullscreen mode Exit fullscreen mode

That printed 200 139532 for me. Three things tripped me up on the way there:

  1. Leave out the Content-Type header and you get a 415. curl's default for -d is form encoding, and the endpoint wants JSON.
  2. Python's default user agent gets a 403. The exact same request from urllib failed until I set any custom User-Agent. A plain release-note-watcher/0.1 was enough. (I checked by sending Python-urllib/3.14 from curl too. 403 again, so it's the UA string, not something else urllib does.)
  3. Every block comes back double wrapped, as {"value": {"value": {...the block...}}}. I peel layers until I hit a dict with a type key, so the code doesn't break if they ever drop one level.

The page has 219 top-level children and limit: 100 returned 104 blocks, so you don't get everything in one call. For a changelog that's fine because the newest entries are at the top. If you need the full history there's a cursor in the response you can feed back in, but I haven't needed it.

One more surprise: the entries aren't Notion headings. Each one is a plain text block like "September 13, 2026 (2026-09-13-base, Model 4.10b aka 4o) | released on September 18, 2026", followed by bulleted items. So I match the date pattern instead of looking for header blocks:

import json
import re
import urllib.request

PAGE_ID = "6f58686f-6381-498b-aef3-5212463b7da6"
API = "https://gptzero.notion.site/api/v3/loadPageChunk"
ENTRY = re.compile(r"^[A-Z][a-z]+ \d{1,2}, \d{4} \(")  # "September 13, 2026 (..."


def load_chunk(page_id, limit=100):
    payload = {"pageId": page_id, "limit": limit, "cursor": {"stack": []},
               "chunkNumber": 0, "verticalColumns": False}
    req = urllib.request.Request(
        API,
        data=json.dumps(payload).encode(),
        headers={"Content-Type": "application/json",
                 "User-Agent": "release-note-watcher/0.1"},  # default UA gets a 403
    )
    with urllib.request.urlopen(req, timeout=30) as resp:
        return json.load(resp)


def unwrap(record):
    while isinstance(record, dict) and "value" in record and "type" not in record:
        record = record["value"]
    return record


def text(block):
    return "".join(seg[0] for seg in block.get("properties", {}).get("title", []))


blocks = {k: unwrap(v) for k, v in load_chunk(PAGE_ID)["recordMap"]["block"].items()}
entries = 0
for child_id in blocks[PAGE_ID]["content"]:
    block = blocks.get(child_id)
    if not block or "type" not in block:
        break  # past what this chunk returned
    line = text(block).strip()
    if block["type"] == "text" and ENTRY.match(line):
        entries += 1
        if entries > 3:
            break
        print("\n" + line)
    elif entries and block["type"] == "bulleted_list" and line:
        print("  - " + line)
Enter fullscreen mode Exit fullscreen mode

Output when I ran it on 2 October 2026:

September 13, 2026 (2026-09-13-base, Model 4.10b aka 4o) | released on September 18, 2026
  - Decreased false positive rate on a variety of domains
  - Improved performance on third party benchmarks

August 12, 2026 (2026-08-09-base, Model 4.9b) | released on August 12, 2026
  - Improved recall on frontier LLMs including Claude 5, GPT 5.6, Gemini 3.6, and Grok 4.5 series models

August 1, 2026 (2026-08-01-base, Model 4.8b) | released on August 1, 2026
  - Robustness to headers; headers are now excluded from our model’s input
Enter fullscreen mode Exit fullscreen mode

Fair warning: loadPageChunk is Notion's internal API, not a documented one. It works today and could change without notice. One request per run, and I only run it once a day.

The page record also has a last_edited_time. I print it, but I don't trust it as a change signal yet, because I haven't watched it long enough to know whether editing a child block bumps the page's timestamp. Hashing the text is cheap, so that's what I compare.

Problem 2: a help center behind Cloudflare, and a timestamp that lies

Turnitin's documentation lives at guides.turnitin.com. curl on any article page gets a 403 and a 5.6 KB page titled "Just a moment...". But the site is a Zendesk Help Center, and Zendesk's public Help Center API answers the same article as JSON with no fuss:

curl -s https://guides.turnitin.com/api/v2/help_center/en-us/articles/22774058814093.json \
  | python3 -c "import json,sys; a=json.load(sys.stdin)['article']; print(a['title'], a['edited_at'], a['updated_at'], len(a['body']))"
Enter fullscreen mode Exit fullscreen mode

You get the full HTML body plus timestamps. And here's the part that made me rewrite my first version of this.

There are two timestamps, and the obvious one is the wrong one. updated_at sounds like "the article changed". It isn't. On the same article (it's called "Using the AI Writing Report") I pulled it around 03:44 UTC and again around 04:54 UTC. In between, updated_at went from 2026-10-02T02:57:47Z to 2026-10-02T04:23:31Z. Same body, byte for byte: 9,671 characters, same hash. Meanwhile edited_at sat at 2026-08-26T09:37:07Z both times.

It isn't one weird article either. I pulled the 100 most recently updated articles (sort_by=updated_at&sort_order=desc) at about 04:55 UTC. 21 of them had an updated_at from that same day. Zero had an edited_at from that day. The page called "Welcome to Turnitin Guides" was last edited in July 2024, and its updated_at said 04:33 UTC that morning.

Zendesk's own Help Center API reference is pretty clear once you read it closely. edited_at is "The time the article was last edited in its displayed locale", and in the sorting section, "order by the last time the title or body was edited". updated_at is just "The time the article was last updated". I don't know what keeps touching updated_at, and the docs don't say. I just stopped using it.

Bonus: sort_by=edited_at works on the list endpoint, which gives you a real "recently changed" feed:

curl -s "https://guides.turnitin.com/api/v2/help_center/en-us/articles.json?per_page=10&sort_by=edited_at&sort_order=desc" \
  | python3 -c "import json,sys; [print(a['edited_at'], a['title']) for a in json.load(sys.stdin)['articles']]"
Enter fullscreen mode Exit fullscreen mode

The locale is already in the path (en-us), which the docs say edited_at sorting needs.

The whole watcher

Both sources go into one small script that keeps a state.json next to it and prints what moved since last time. Standard library only:

import hashlib
import json
import pathlib
import urllib.request

UA = {"User-Agent": "release-note-watcher/0.1"}
STATE = pathlib.Path("state.json")
ZD = "https://guides.turnitin.com/api/v2/help_center/en-us/articles/{}.json"
TURNITIN_IDS = [28294949544717, 22774058814093]


def get_json(url, payload=None):
    headers = dict(UA)
    data = None
    if payload is not None:
        headers["Content-Type"] = "application/json"
        data = json.dumps(payload).encode()
    req = urllib.request.Request(url, data=data, headers=headers)
    with urllib.request.urlopen(req, timeout=30) as resp:
        return json.load(resp)


def turnitin(article_id):
    a = get_json(ZD.format(article_id))["article"]
    body_hash = hashlib.sha256(a["body"].encode()).hexdigest()[:12]
    return {"title": a["title"], "edited_at": a["edited_at"], "body": body_hash}


def gptzero_entries():
    page_id = "6f58686f-6381-498b-aef3-5212463b7da6"
    data = get_json(
        "https://gptzero.notion.site/api/v3/loadPageChunk",
        {"pageId": page_id, "limit": 100, "cursor": {"stack": []},
         "chunkNumber": 0, "verticalColumns": False},
    )
    blocks = {}
    for key, rec in data["recordMap"]["block"].items():
        while isinstance(rec, dict) and "value" in rec and "type" not in rec:
            rec = rec["value"]
        blocks[key] = rec
    lines = []
    for child in blocks[page_id]["content"]:
        b = blocks.get(child)
        if not b or "type" not in b:
            break
        title = b.get("properties", {}).get("title", [])
        lines.append("".join(seg[0] for seg in title))
    return hashlib.sha256("\n".join(lines).encode()).hexdigest()[:12]


now = {"gptzero": gptzero_entries()}
for aid in TURNITIN_IDS:
    now[f"turnitin:{aid}"] = turnitin(aid)

old = json.loads(STATE.read_text()) if STATE.exists() else {}
for key, value in now.items():
    if key not in old:
        print(f"new    {key}: {value}")
    elif old[key] != value:
        print(f"CHANGE {key}: {old[key]} -> {value}")
    else:
        print(f"same   {key}")
STATE.write_text(json.dumps(now, indent=2))
Enter fullscreen mode Exit fullscreen mode

First run:

new    gptzero: 66ddc34ba254
new    turnitin:28294949544717: {'title': 'AI writing detection model', 'edited_at': '2026-05-15T08:03:20Z', 'body': 'fbebac616689'}
new    turnitin:22774058814093: {'title': 'Using the AI Writing Report', 'edited_at': '2026-08-26T09:37:07Z', 'body': '69f666c6c94a'}
Enter fullscreen mode Exit fullscreen mode

Second run, right after: three lines of same. To make sure the change path actually works and isn't just always printing same, I hand-edited one edited_at in state.json and ran it again. It printed a CHANGE line for that article and same for the other two.

The first Turnitin ID is their model changelog page. Its updated_at that morning said 30 September. Its edited_at said 15 May. If I'd been watching updated_at I would have gone and reread an article that hadn't changed in four and a half months.

What this doesn't tell you

A release note is the vendor's description, not a log of what ran on a given day. GPTZero's own notes say the new model was "released on September 18, 2026", and their blog post about the same model says it became the default for all users on September 20. If you care about one specific scan from September 19, the notes alone won't settle it.

So my text file now has two columns per change, "what the notes say" and "what I could confirm", and the second column is often empty. That still beats guessing.

If anyone knows a cleaner way to read public Notion pages than the internal endpoint, I'd like to hear it. That's the piece I trust least.

Top comments (0)