DEV Community

Cover image for Google Cloud Skill Registry: A Practical Guide for Engineers
Karthigayan Devan
Karthigayan Devan

Posted on

Google Cloud Skill Registry: A Practical Guide for Engineers

Introduction

If you have built more than two or three AI agents, you have probably felt this pain already. Every agent needs "know-how": how to query BigQuery the right way, how your team names GKE clusters, how to file a ticket, how to run a cost report. And every time, someone copies the same prompt snippets and helper scripts into yet another repo.

A few months later you have five slightly different versions of the same instructions, nobody knows which one is current, and the agent's context window is stuffed with things it does not need for the task in front of it.

Google Cloud's Skill Registry (part of Gemini Enterprise Agent Platform, currently in Preview) is aimed right at this problem. Think of it as a private package registry, but for agent skills instead of npm or pip packages. You publish a skill once, it gets versioned, and agents can search for it and load it only when the user's request actually needs it.

In this post I will walk through what Skill Registry is, how the pieces fit together, what a skill package looks like, and the full REST API flow with working curl commands. I will also call out the limits and gotchas I noticed, so you know what you are signing up for before you build on a Preview service.

What Skill Registry actually is

Skill Registry is a secure, private, low-latency store for agent skills. A skill is a self-contained package: instructions in a SKILL.md file, plus any scripts, reference docs, and assets the agent needs to do one job well.

The API is built around just two resources, and once you get these, the rest of the API makes sense.

Resource Mutable? What it holds
Skill Yes Metadata (display name, labels, create and update time), the default revision, and the skill content
Skill revision No An immutable snapshot of one version: name, description, and a fixed pointer back to its parent skill

So the Skill is the thing you keep editing, and every meaningful change leaves behind a revision you can go back and inspect. If you have worked with Cloud Run services and revisions, this model will feel very familiar.

There is also one skill you get for free: gcp-skill-registry. It is a built-in skill, managed and versioned by Google, that teaches an agent how to talk to Skill Registry itself. With it, an agent can search for skills, create new ones, and manage what is available. One small surprise: Google provisions it on your first API call, so that first response will not list it. It shows up from the second call onward.

How it fits together

There are really two paths through Skill Registry: a publish path for the people writing skills, and a retrieval path for the agents using them.

  PUBLISH PATH                     SKILL REGISTRY (one region)             RETRIEVAL PATH

+----------------------+        +-----------------------------------+
| Skill author or CI   |        |  Skills API (v1beta1)             |
| zip SKILL.md,        |------->|  create, update, delete, list,get |
| scripts, references, |        +-----------------------------------+
| assets + base64      |                         |
+----------------------+                         v
          |                     +-----------------------------------+
          | polls the           |  Async validation (LRO)           |
          +-------------------->|  zip safety + SKILL.md checks     |
            operation           |  fails the operation, not POST    |
                                +-----------------------------------+
                                                 |
                                                 v
                                +-----------------------------------+
                                |  Skills + immutable revisions     |
                                |  incl. built-in gcp-skill-registry|
                                +-----------------------------------+
                                                 |
                                                 v                     +----------------------+
                                +-----------------------------------+  | AI agent             |
                                |  RetrieveSkills (semantic search) |<-| ADK / Managed Agents |
                                |  matches display name+description |->| searches by intent,  |
                                +-----------------------------------+  | loads only what fits |
                                                                       +----------------------+
Enter fullscreen mode Exit fullscreen mode

On the publish path, you send a zipped, base64-encoded skill to the API. Validation runs in the background as a long-running operation, and only a valid package becomes a skill with a new immutable revision. On the retrieval path, an agent (built with ADK or the Managed Agents API) describes what it needs in plain words, RetrieveSkills does a semantic match on each skill's display name and description, and the agent loads just the skills that fit. When a skill is attached to an agent, its SKILL_ID becomes the folder name the agent sees.

Anatomy of a skill

A skill is just a folder that you zip up. The only file that must be there is SKILL.md. Everything else is optional, but the docs package a typical skill like this:

finops-cost-report/
├── SKILL.md          # required: front matter + instructions
├── scripts/          # helper code the agent can run
│   └── monthly_cost.py
├── references/       # docs the agent can read when needed
│   └── billing_export_schema.md
└── assets/           # templates, sample files, etc.
    └── report_template.md
Enter fullscreen mode Exit fullscreen mode

SKILL.md starts with YAML front matter and then plain Markdown instructions. Here is a small example:

---
name: finops-cost-report
description: Builds a monthly Google Cloud cost report per project from the BigQuery billing export. Use when someone asks for spend, cost trends, or top cost drivers.
license: Apache-2.0
---

# FinOps cost report

1. Read references/billing_export_schema.md to understand the table.
2. Run scripts/monthly_cost.py with the project ID and month.
3. Fill assets/report_template.md with the results.
4. Call out any service whose cost grew more than 20% month over month.
Enter fullscreen mode Exit fullscreen mode

The front matter has strict rules, and the registry checks them:

  • name: required, up to 64 characters, lowercase letters, numbers, and hyphens only, and it cannot start or end with a hyphen.
  • description: required, up to 1,024 characters. Write this one carefully, because semantic search uses it to decide if your skill matches a request.
  • license: optional, up to 1,024 characters.
  • The instructions body: up to 500,000 characters.

A tip from experience with any kind of retrieval: write the description the way a user would describe their problem, not the way you would describe your code. "Use when someone asks for spend or cost trends" will match far better than "Runs a SQL aggregation on billing data."

For real-world examples, Google keeps sample SKILL.md files in the Google Cloud Skills repository.

Hands-on: the full lifecycle with the REST API

The docs also show Python and Node.js samples, plus an "Intro to Skill Registry" notebook you can open in Colab. I am sticking with curl here because it shows exactly what goes over the wire.

Step 0: Set up your project

  1. Pick or create a Google Cloud project and make sure billing is on.
  2. Enable the Agent Platform API (aiplatform.googleapis.com).
  3. Grant yourself roles/aiplatform.user (or roles/aiplatform.viewer for read-only) plus roles/serviceusage.serviceUsageConsumer. Skill Registry inherits project-level IAM, so there is no separate permission model to learn.

Then set a few variables so the rest of the commands stay readable:

export PROJECT_ID="my-agent-project"
export LOCATION="us-central1"   # or europe-west4, us-east5
export SKILL_ID="finops-cost-report"
export BASE="https://${LOCATION}-aiplatform.googleapis.com/v1beta1/projects/${PROJECT_ID}/locations/${LOCATION}"
export TOKEN="$(gcloud auth application-default print-access-token)"
Enter fullscreen mode Exit fullscreen mode

Step 1: Package the skill

The API expects the whole skill folder as a zip file, encoded as one single-line base64 string.

cd finops-cost-report
zip -r skill.zip SKILL.md scripts/ references/ assets/
base64 -w 0 skill.zip > skill.b64   # on macOS: base64 -i skill.zip -o skill.b64
Enter fullscreen mode Exit fullscreen mode

Step 2: Create the skill

A zipped skill can be up to 10 MB, which is far too long to paste inline into a shell command. So I build the JSON body in a file first and point curl at it.

cat > create.json <<EOF
{
  "displayName": "finops-cost-report",
  "description": "Builds a monthly Google Cloud cost report from the BigQuery billing export.",
  "zippedFilesystem": "$(cat skill.b64)"
}
EOF

curl -X POST \
  -H "Authorization: Bearer ${TOKEN}" \
  -H "Content-Type: application/json" \
  -d @create.json \
  "${BASE}/skills?skillId=${SKILL_ID}"
Enter fullscreen mode Exit fullscreen mode

Pick your SKILL_ID with care. It must be 1 to 63 characters, lowercase letters, numbers, and hyphens, start with a letter, and it cannot start with gcp- (Google reserves that prefix). It is also permanent: once used, it stays reserved even after you delete the skill. And since the SKILL_ID becomes the folder name when the skill is attached to an agent, a descriptive name actually helps the model understand what the skill is for.

Step 3: Wait for the long-running operation

Create, update, and delete do not finish right away. They return an operation, and validation of your zip happens in the background. The response contains a name like projects/PROJECT_NUMBER/locations/LOCATION/skills/SKILL_ID/operations/OPERATION_ID. Poll it until done is true:

OP_NAME="projects/123/locations/us-central1/skills/finops-cost-report/operations/456"

curl -s -H "Authorization: Bearer ${TOKEN}" \
  "https://${LOCATION}-aiplatform.googleapis.com/v1beta1/${OP_NAME}"
Enter fullscreen mode Exit fullscreen mode

This step matters more than it looks. A bad zip does not fail on the POST. The POST succeeds, and the operation fails later with a validation error. If your CI pipeline only checks the HTTP status of the create call, it will happily report green on a broken skill.

Step 4: Update the skill

Update is a PATCH, and only the fields you list in updateMask get changed. Leave zippedFilesystem out if you only want to fix the description.

curl -X PATCH \
  -H "Authorization: Bearer ${TOKEN}" \
  -H "Content-Type: application/json; charset=utf-8" \
  -d '{"description": "Monthly cost report per project, with top cost drivers and month-over-month growth."}' \
  "${BASE}/skills/${SKILL_ID}?updateMask=description"
Enter fullscreen mode Exit fullscreen mode

Step 5: List, get, and inspect revisions

# All skills in this region (name, display name, description, state)
curl -s -H "Authorization: Bearer ${TOKEN}" "${BASE}/skills"

# One skill: metadata plus the payload of its latest revision
curl -s -H "Authorization: Bearer ${TOKEN}" "${BASE}/skills/${SKILL_ID}"

# Revision history
curl -s -H "Authorization: Bearer ${TOKEN}" "${BASE}/skills/${SKILL_ID}/revisions"

# One specific revision (ID is the last part of the revision name)
curl -s -H "Authorization: Bearer ${TOKEN}" "${BASE}/skills/${SKILL_ID}/revisions/4567890123"
Enter fullscreen mode Exit fullscreen mode

A list call returns entries like this:

{
  "name": "projects/1234567890/locations/us-central1/skills/3456789012",
  "createTime": "2026-05-10T00:02:12.497720Z",
  "updateTime": "2026-05-10T00:02:19.064874Z",
  "displayName": "cymbal_skill",
  "description": "A skill for managing Cymbal projects.",
  "state": "ACTIVE"
}
Enter fullscreen mode Exit fullscreen mode

Step 6: Find skills with semantic search

This is the most interesting call. RetrieveSkills takes a plain-language query and matches it against each skill's display name and description. This is what lets an agent pick the right skill at runtime instead of loading everything up front.

curl -s -G -H "Authorization: Bearer ${TOKEN}" \
  --data-urlencode "query=find skills to report on cloud spend" \
  "${BASE}/skills:retrieve"
Enter fullscreen mode Exit fullscreen mode

Step 7: Delete a skill

Delete removes the skill and all of its revisions, and it is also a long-running operation. You cannot delete the built-in gcp- skills.

curl -X DELETE -H "Authorization: Bearer ${TOKEN}" "${BASE}/skills/${SKILL_ID}"
Enter fullscreen mode Exit fullscreen mode

Note: Always double check request paths against the live docs before you script anything important, especially the retrieve path, since this is a Preview API and paths can change.

Validation rules cheat sheet

Every create or update runs your zip through a set of safety checks. Most of them exist to block classic zip attacks (zip bombs, path traversal, symlink tricks), which is reassuring when agents are going to run code from these packages.

Check Limit or rule
Zipped archive size 10 MB max
Total unzipped size 500 MB max
Number of items in the zip 10,000 max
Compression ratio 100 max
Folder depth 8 levels max
File and folder names No .., no leading / or \, no duplicates
Symbolic links Not allowed
Empty zip or invalid zip Rejected
SKILL.md Must exist, with YAML front matter and Markdown content
name (front matter) Required, 64 chars max, lowercase, numbers, hyphens, no leading or trailing hyphen
description (front matter) Required, 1,024 chars max
license (front matter) 1,024 chars max
Instructions body 500,000 chars max

My suggestion: copy these rules into a small pre-flight script in your CI pipeline. Checking for SKILL.md, the name format, and the zip size locally takes a few lines of Python and saves you a round trip to a failed long-running operation.

Regions, compliance, and the gotchas

Skill Registry runs in three regions today: us-central1 (Iowa), europe-west4 (Netherlands), and us-east5 (Columbus, Ohio). Skills live in a region, so if your agents run in another region, plan for that.

Compliance feature Status
Access Transparency Supported
Data Residency in US and EU Supported
Customer-Managed Encryption Keys (CMEK) Not supported
HIPAA Not supported
VPC Service Controls Not supported

Before you put this in front of production agents, keep these points in mind:

  • It is Preview. Pre-GA terms apply, support is limited, and the API is v1beta1. Expect changes.
  • No VPC-SC, CMEK, or HIPAA yet. For regulated workloads (healthcare, or anything behind a VPC-SC perimeter), this is a blocker for now.
  • Skill IDs are forever. A SKILL_ID stays reserved even after deletion, and the docs also mention a 24-hour wait before reuse. Either way, do not use throwaway IDs in shared projects.
  • Validation is async. A successful POST does not mean a valid skill. Always poll the operation.
  • The built-in skill appears late. gcp-skill-registry is not in your very first API response.
  • Search quality depends on you. RetrieveSkills only looks at display name and description. A vague description means a skill your agents never find.
  • Skills run code. The registry checks the package structure, not what your scripts actually do. Review skill code the same way you review any code that runs with your agent's credentials.

One more thing worth knowing: Google also lets you govern standalone skills centrally through Agent Registry. If your company is building a lot of agents, look at both and decide where skills should be governed.

Conclusion

Skill Registry takes something most teams are doing badly today, copying prompts and helper scripts between agent repos, and turns it into a proper platform service. You get versioning through immutable revisions, security checks on every package, IAM you already understand, and semantic search so agents load only what they need.

It is still early. The Preview label, three regions, and missing VPC-SC and CMEK support mean I would not move regulated production workloads onto it yet. But for platform teams that want one shared, governed home for agent know-how, it is well worth a weekend of hands-on time now, so you are ready when it goes GA.

My advice: start with one or two skills your team already copies around, write really good descriptions, wire the long-running operation check into CI, and see how well RetrieveSkills picks them up.

Have you tried Skill Registry, or are you managing agent skills some other way? I would love to hear how in the comments.

Resources

Top comments (0)