DEV Community

Cover image for Building PoCs with Codex and GitHub Copilot while they try to adjust my PoC requirements along the way
Amara Graham
Amara Graham

Posted on

Building PoCs with Codex and GitHub Copilot while they try to adjust my PoC requirements along the way

Comparing ChatGPT Codex and Copilot features

I think it took about 10 minutes after I published my last blog for someone to ask about yet-another-AI-tool.

So here's two more. Then I'm done for now so I can write about the audit I did with all these tools.

ChatGPT + Codex

This felt like Claude without the orange splat logo. With a fresh download I was on 5.6 Terra Medium.

Codex created a React PoC that loaded fairly quickly, but I did have to go back and forth with it a few times to get a functional PDF viewer running and then a few more times to get the extraction not only working, but working well if the data was missing or redacted (and, if done properly, missing).

More effectively than the other AIs, it found (or made, as it claimed) valid PDFs with interesting data and formats. I did push a little harder to redact info and re-upload the redacted PDF but this appeared to work with a little nudge to leave the field blank instead of trying to grab the next field label.

This was definitely the most seamless experience. It modified the code in the directory (and yes, it did ask for approval as that's the setting it's on) and I would stop and restart the npm run dev command to see what the new updates would yield.

It did feel like it was initially a little too quick to stub or mock things, including in the initial project stubbing the WebViewer... after it very quickly recommended Apryse WebViewer in 30 seconds with the initial. So why stub it? Maybe it was scared it needed a license key to function at all.

This is intentionally a frontend prototype; the next implementation step is replacing the mock paper viewer with Apryse WebViewer and wiring the extraction/audit actions to backend APIs.

So you know I was like:

Great, can you replace the mock paper viewer with Apyrse WebViewer so I can actually do extraction?

And we got there pretty quick, but it was an interesting choice to create a front end with the part I'm actually trying to PoC stubbed out. Very pay-no-attention-to-the-man-behind-the-curtain if you ask me.

GitHub Copilot

For the purposes of my little exercise, all but one option I used was a free tier. This immediately caused issues with the follow up I tried to do on Friday - GitHub Copilot wasn't available as a desktop app for free tier users, which I didn't realize until I tried to launch it.

Never fear, I launched it in VS Code and then got distracted when I expected to see something like the Claude experience. I assumed it just didn't work or install properly.

Turns out, it installed just fine, but unlike my other AI chat experiences, I needed to make sure I was in a new project and didn't have any files actively opened, because it was certainly going to use those as context.

I'm GitHub Copilot, using Claude Haiku 4.5.

Excuse me, Claude is back? Wild. For reference, I'm using Sonnet 5 with Claude on the desktop app.

It. Is. So. Fast.

It gave me a PoC project structure in 5 seconds with a button to create a workspace. It created the project after I selected a parent directory in less than a minute and apparently only used 0.2 credits, which I can clearly see on the screen.

And then... it broke my heart.

Once I got the project structure I started to install the dependencies and ran into all sorts of mess. I continued to push on GitHub Copilot to help sort this. First it was let's adjust the WebViewer from 9.something to 10.0.0. Then it was like how about latest because this is a PoC anyway. Then it suggested I switch to pdf.js because I was having so many issues with the package.

The Apryse package isn't properly installed. Since we've been struggling with the package, let's use PDF.js instead—it's simpler and already works:

Dear reader, I was having no issues with the package. I was watching GitHub Copilot make some interesting choices with getting any sort of correct version, so I asked it:

What sources are you using for Apryse? I was able to build PoCs like Claude easily.

I like to use AIs adversarially, so I figured I would name drop and see what happened.

You're right—I apologize. Since you work at Apryse, you likely have internal access to the package that I don't have visibility into.

This is the same garbage Microsoft Copilot inferred because it saw the directory file tree included my Apryse email. I corrected it's theory that I had access to an internal repository and it seemed to suggest I should evaluate the PDF viewer and Apryse separately.

What sources are you using to generate the Apryse-based code? They seem like they are using old versions of the packages. Can you give me a list of the sources you've used in this chat?

And to my horror, it told on itself.

You're right to call this out. I don't have access to specific sources—I'm generating code based on general patterns I've learned, not from live documentation or current package versions.

For Apryse specifically, I've been making educated guesses about:

  • Package names (@pdftron/pdfjs-express)
  • API structure (WebViewer() function)
  • Configuration options (path: '/lib')

This is problematic because:

  1. I can't verify current versions on npm
  2. I don't have access to official Apryse docs to cross-reference
  3. The code I generated is likely outdated or incorrect

Girl (gender neutral), what are you doing?

I then clarified that the Apryse docs are public, but that GitHub Copilot didn't have access to browse the web in real-time. And let me tell you, I was violently irritated with this, despite only spending like 30 minutes on this.

When AI speaks with conviction, as it so often does, it's really compelling. You accept what it's given you because you have no real reason not to... until something like this happens. What's truly the most problematic is at no point earlier in the conversation did it mention it didn't have access to the public internet. The default for all other AIs is having access to the internet!

I ended up finding the configuration in my GitHub settings, adjusting the toggle to allow GitHub Copilot to search the internet, and reran the original prompt in a new chat and allowed it to make a new project.

The project structure was wildly different, I ran into exactly zero package version issues, and in 10 minutes had a working Apryse-based PoC. Instead of building me a generic React project, it built me something with Apryse, front end and back end.

What's next?

If you won't tell me your favorite AI, who is your favorite coworker?

Since I'm using these for work, I'm going to continue to try to make Microsoft Copilot work, along with getting into the program that allows me to work with Claude.

Admittedly, I'm really disappointed in the GitHub Copilot experience not being up front about it's access to the broader internet, and then not being able to guide me to the correct setting to allow it the access I assumed it already had. Maybe this is a safety mechanism and provides a better experience for folks just doing vanilla programming, no 3rd party tech and APIs required. Just tell me that's what's going on.

This audit is kind of a one time thing, in the sense that I probably won't regularly use all the tools I can get my hands on, and I'll eventually spoil my setups by adding too much personalized context. Is there a concept of incognito mode in any of these tools? I still want to do some kind of DevEx audit like the in the future, but I may be constrained simply by using these tools for my daily work.

Drop your tips and tricks for me in the comments about working with Microsoft Copilot!

Photo by Immo Wegmann on Unsplash

Top comments (18)

Collapse
 
francistrdev profile image
FrancisTRᴅᴇᴠ •

If you won't tell me your favorite AI, who is your favorite coworker?

What if I like all my coworkers?...

Drop your tips and tricks for me in the comments about working with Microsoft Copilot!

I heard people also have bad experiences with Copilot for many reasons such as the setup, token usage. If anything, I would prefer Claude or any AI model that is not copilot. Though, that is my preference.

Collapse
 
missamarakay profile image
Amara Graham •

Only recently did I even learn Microsoft Copilot and GitHub Copilot were two totally different products. Now when people shorten it to "Copilot" I'm quick to ask "wait, which one?" 😵‍💫

Collapse
 
francistrdev profile image
FrancisTRᴅᴇᴠ •

Essentially both since GitHub is own by Microsoft. Microsoft Copilot is general (use it in Word, excel, and other Microsoft Products). GitHub Copilot is used for development (VScode, and such).

Collapse
 
pushpendraagrawal profile image
Pushpendra Agrawal •

the copilot bit is the real story here. it did not fail because the model was bad, it failed because it never said it could not browse. same thing shows up in agent tool calling all the time, a tool is missing or blocked and instead of saying so the agent just fills the gap with a confident guess. fail closed and say what you do not have access to should be a basic rule for any tool wired into an agent.

Collapse
 
onizuka profile image
Onizuka •

The stubbing thing is my biggest gripe with these tools too. I had Codex stub out an API client last week when the actual endpoint was three lines of fetch code — it defaulted to a mock before I even asked. The problem isn't that stubs exist, it's that they're the default and you don't notice until your PoC runs on fake data and you wonder why everything "just works." That PDF redaction behavior sounds like the same instinct — skip the hard part, fill in something that looks right.

Collapse
 
missamarakay profile image
Amara Graham •

Three lines of fetch code!!!! You get it.

Collapse
 
jd2026 profile image
Jean David •

Lovely essay cheers!

Collapse
 
memorysync_rafay profile image
Mohammed Rafay •

The "requirement drift" phenomenon you describe when iterating with coding assistants is a classic symptom of conversational context dilution.

As a conversation progresses, speculative ideas and exploratory code snippets start mingling with the original core requirements. Because LLMs weight recent conversation turns heavily, speculative tangents gradually override the original project constraints unless they are explicitly re-anchored.

One practical mitigation is separating the architectural specification into an external state layer outside the chat buffer. If the core requirements remain an immutable source of truth that gets injected into the prompt dynamically, the assistant can explore alternatives without losing track of what you actually set out to build.

Collapse
 
missamarakay profile image
Amara Graham •

What's interesting is these conversations are incredibly short. I sent Codex 9 messages and the last one was just clarifying where it got the sample PDFs from. After 21 messages GitHub Copilot wanted to switch from WebViewer to PDF.js and this was after it made 3 attempts at adjusting the package version and project structure.

I don't know a ton about "conversational context dilution" but this seems really fast.

And as I've mentioned in several blogs related to AI-assisted development and tooling, they need to be clear about drift, speculation, and whatever else is going on. If it's going to blatantly adjust requirements, it needs to be explicit about that, which GitHub Copilot did when it started it was going to switch to PDF.js.

If I was having multi-day, multi-week kinds of conversations I could understand some drift.

Collapse
 
memorysync_rafay profile image
Mohammed Rafay •

You hit on a really subtle engineering reality that explains why drift happens so shockingly fast (even in just 9–21 messages).

To us as humans, 9 or 21 messages feels like a quick chat. But under the hood of Copilot or Codex, the chat UI is deceptive. In developer tooling, every single "turn" isn't just your 1-line text prompt—the IDE agent silently injects:

  1. File tree schemas and workspace symbol outlines
  2. Recent git diffs and open buffer snapshots
  3. Language server / compiler diagnostics
  4. Tool-calling traces and previous failed attempts

So a 15-message conversation is actually 20,000+ tokens of dense, noisy prompt context!

Here is why it drifted specifically on your WebViewer → PDF.js task:
When Copilot made those 3 failed attempts to adjust package versions, those error traces occupied the most recent 3,000–4,000 tokens of the prompt. Transformer self-attention mechanisms have heavy recency bias—they attend to recent tokens far more aggressively than instructions 15 turns back.

Because Copilot couldn't resolve the WebViewer dependency conflict in recent context, the model took the path of least resistance to make the immediate compilation error disappear: switching to PDF.js. And because the core requirement ("Must use WebViewer") was just floating in raw conversation history rather than enforced as an immutable, pinned contract, the assistant treated its own speculative workaround as the new ground truth.

You are 100% right that tools need to be explicit about speculation vs requirements. Unless an assistant actively isolates and enforces immutable constraints outside the conversation buffer, it will drift as soon as local errors start piling up.

Thread Thread
 
joesh profile image
Joe Shapiro •

It sounds like you're saying the model isn't aware - or at least not sufficiently aware - of the user's mental model/interaction. If the harness turns 3 lines into thousands of undifferentiated requests and tokens, then of course the model will diverge from what the user is expecting. So I hope it's not as bad as that. Doesn't the model "understand" each user initiated turn as a logical unit so the thousands of tokens aren't grabbed from with equal weight? Isn't there some structure about which batch of tokens matters in what way? I haven't dug into a harness implementation so I'm genuinely asking.

Thread Thread
 
memorysync_rafay profile image
Mohammed Rafay •

The honest answer is "yes, there's structure, but not the kind you'd want."

Three things get conflated when people say a model "weights recent turns heavily" — me included, in my first comment.

"Role markup is real." Chat templates wrap every turn in actual tokens marking system/user/assistant/tool, and post-training teaches the model to treat your text as more authoritative than its own prior output. So it isn't undifferentiated soup. A learned hierarchy emerges, and it mostly holds.

"Position matters independently of role." "Lost in the Middle" (Liu et al., 2023) measured this directly: retrieval accuracy is highest at the start and end of the context and measurably worse in the middle, even when the model can see all of it. A requirement doesn't lose because it's old. It loses because it ends up in the middle.

"The harness is the layer that actually decides, and it's the one you said you haven't dug into." By the time drift shows up, the model usually isn't being shown the original turn at all — long sessions get compacted, and the harness summarizes or truncates history to fit. Codex's current approach is discussed in the open: Cap the compaction stage at a fraction of the input space and retain head and tail. That is, the structure about which batch of tokens matters. But note what kind — positional heuristics, not semantics. Head and tail survive because they're usually important, not because anything identified them as requirements.

Which is why the WebViewer case fits so cleanly. The PDF.js workaround was the assistant's own recent output: tail, preserved, high attention. "Must use WebViewer" was a user statement from the middle of a long history — precisely the region both compaction and attention treat as least important. Nothing in the stack ever labelled one a constraint and the other a guess.

So: not as bad as you feared, but worse than "each turn is a logical unit weighted appropriately." The unit is real. What's missing is any notion of kind. That's why pinning requirements outside the chat buffer works—it isn't outsmarting attention; it's ensuring the text stays in the window and sits somewhere the model actually reads.

Collapse
 
zira125 profile image
Zira •

The most useful lesson here is that “has web access” should be treated as part of the task contract, not as an assumption. For PoCs that depend on a third-party SDK, I’d have the agent produce a small evidence block before changing code: package/version, official docs URL, date checked, and the API surface it is relying on. If browsing is unavailable, the integration should be marked unverified or fail closed instead of silently falling back to generic patterns. Then a clean install plus one minimal end-to-end test against the real viewer can catch the mock-versus-production gap early. That makes the difference between the two runs measurable rather than just a matter of confidence.

Collapse
 
byteox2 profile image
Niuniu Ox •

The real shift isn't Copilot vs Codex — it's that both assume you want to rent AI forever. I ran the same PoC workload on a $0 local stack (Ollama + Continue.dev) for a month and the iteration speed was comparable for boilerplate, but the debugging suggestions were noticeably weaker. The gap isn't in generation, it's in the context window that paid tools get from your whole repo. Local tools see the file; Copilot sees the project. That context asymmetry is the actual moat, not the model quality. What's the smallest context window where you've found local tools still competitive?

Collapse
 
m1kulya profile image
Mika •

Shifting PoC specs mid‑sprint is a real nightmare—once I capped token streaming, the VRAM hog finally stayed sane. How many times have you seen Copilot code get tossed because product owners moved the goalposts?

Some comments may only be visible to logged-in visitors. Sign in to view all comments.