DEV Community

Cover image for Our support agent recommended replacing a valid API key
Pierre-Laurent Medori
Pierre-Laurent Medori

Posted on

Our support agent recommended replacing a valid API key

On September 22, our support diagnostic produced the same wrong recommendation three times for one customer's app: replace its OpenAI key.

The key authenticated. The customer's project could not access the model we used to test it.

The recommendation was already written in our Python code:

Check the OpenAI secret key configured on the section and replace it.

The model received a verdict and a next step. We had written both.

I'm currently supervising the support agent at GoodBarber, a SaaS platform for building apps. It uses internal diagnostics to inspect the configuration relevant to a customer's question. A diagnostic returns findings, unknowns and recommended actions; the model uses that report to prepare its answer.

Two fixes, on September 21 and 23, exposed how badly we were translating observations into those reports.

A useful answer labelled unavailable

The first investigation covered eleven runs of the ChatGPT integration diagnostic recorded since September 17. All eleven had returned UNAVAILABLE.

We checked the five apps involved. Four had no ChatGPT section, the component where an owner adds the feature and configures its key. Ten runs concerned those apps. The remaining run involved the app with a section and exceeded the diagnostic's eight-second budget.

The absence was useful information. Someone asking why the feature does not work needs to know it has not been configured. But our report labelled that finding UNAVAILABLE, which the prompt defined as a check that could not conclude. The model was told not to assume either way.

We had found the answer and attached a status that prevented it from being used as one.

The fix made confirmed absence an actionable finding: add the section, then configure its key. It also preserved a distinction the diagnostic had been discarding:

None  # The upstream lookup failed.
[]    # The lookup found no ChatGPT sections.
Enter fullscreen mode Exit fullscreen mode

Both had become “no ChatGPT section on this app.” For None, we were describing an app we had failed to inspect. That case now remains explicitly unknown.

The same rule applies when a key check times out: a key is present, but its validity was not verified. It does not count as a failing key.

The model in the error was ours

The September 21 fix made the diagnostic usable. On September 22, it exposed the next defect: the three wrong recommendations that opened this story.

The error code was model_not_found. The accompanying message said the customer's OpenAI project did not have access to gpt-4.1-mini.

That model was hardcoded in our probe; the probe did not read the section's configured model.

To check a key, the tool made a small generation request. It tested authentication, access to our selected model and completion of the request in one operation.

OpenAI documents incorrect keys as a 401 error; here, the response identified the customer's project and denied it access to a model. That project-specific message is the evidence behind our conclusion that the key authenticated.

The error fell through to our default next step: replace the key. Even if the section used the same model, that would call for investigating model access, not automatically replacing the credential.

A GET /v1/models would avoid selecting a generation model, but it requires its own model-listing permission. A successful listing establishes authenticated access to that endpoint; a denied listing still needs interpretation. Our existing probe was a generation check labelled as a key check.

The September 23 correction removes the replacement recommendation for model_not_found and model_not_available. The report explicitly says “the OpenAI key authenticates” and names the inaccessible probe model.

That sentence currently lives in unknowns. If no other section passes or fails conclusively, the overall status remains UNAVAILABLE: the generation check is incomplete. The authentication observation is preserved, but mixing it with the unknown outcome is still awkward. Giving authentication its own finding would make the contract clearer. This fix stops the wrong recommendation; it does not finish that separation.

What the tests have to preserve

The regression tests preserve both outcomes: an inaccessible probe model produces no replacement advice; an explicit invalid_api_key error still produces the corrective step. Suppressing every error would make the diagnostic useless in a different way.

These are specific cases found during the investigation, not a measured failure rate for the whole assistant.

In How do you debug something that is allowed to be wrong?, I traced an agent's repeated writes to a misleading cached read. Here, the tool went further: it turned incomplete evidence into a recommended action.

The prompt tells the assistant to trust live diagnostics over generic guidance. That gives every verdict in the report considerable weight. Before asking whether the model followed the instruction, I need to ask what evidence permitted us to write that verdict.

We had already written “replace it” before the model started answering. The correction belonged where those words were chosen.

When your diagnostic says a credential is invalid, did it test the credential, or its own probe?

Top comments (0)