DEV Community

Михаил
Михаил

Posted on • Originally published at agentlabjournal.online

Connected Apps in Google Search: Testing Three Practical Scenarios

Connected Apps in Google Search: Testing Three Practical Scenarios | Agent Lab Journal

  AL
  Agent Lab Journal


  ← Guides
  Glossary
Enter fullscreen mode Exit fullscreen mode

PRACTICAL TEST · BEGINNER

Connected Apps in Google Search: Testing Three Practical Scenarios

      Level: From scratch
      Reading: 35 minutes
      Result: A comparison table for Instacart, Canva, and YouTube Music
Enter fullscreen mode Exit fullscreen mode

Connecting an app does not prove that the task has become automatic. A useful test must count how many manual actions AI Mode actually removes, measure how long the complete task takes, and identify the exact point where the AI agent stops and waits for a person to choose, save, open, play, publish, or pay.

What this article does—and does not—claim

This is a repeatable test protocol, not a table of invented benchmark results. Connected-app availability, interface wording, supported operations, account requirements, and handoff behavior can differ by region, language, device, account, and rollout stage. Record what your interface actually does.

If one of the three apps is unavailable, that is an availability result, not proof that the app is slow or ineffective. Use the code U, preserve the request you used, and do not replace the missing observation with an estimate.

The test deliberately stops before a real purchase, public design publication, or sharing media with another person. It measures preparation and handoff while keeping consequential external actions under human control.

Contents

  • The question being tested

  • One concrete case for all three apps

  • How to count steps, time, and confirmations

  • Preparing the accounts and observation sheet

  • Running a manual baseline

  • Instacart scenario

  • Canva scenario

  • YouTube Music scenario

  • Repeating the test after connection

  • Comparison tables

  • Verifying the result

  • Failure cases

  • Limitations

  • Drawing a defensible conclusion

The question being tested

Requests such as “build a grocery cart,” “create a flyer,” or “make a playlist” sound like single operations. Each request actually contains a chain of separate decisions:

  • interpret the user’s goal;

  • identify missing information;

  • prepare the requested content;

  • select a third-party application;

  • request access to an account;

  • create or propose an object in that application;

  • hand control to the user;

  • perform or refuse a consequential final action.

The important measurement is the workflow depth reached without manual intervention. A polished response inside Search is not equivalent to a saved cart, editable design, or playable playlist in the target service.

The experiment answers four questions for each app:

  • How many manual actions are required during the first run?

  • How many actions remain after the account is already connected?

  • How much time passes before a verifiable result exists in the target app?

  • Which action does the system leave for the user to confirm or complete?

Keep the first run and the repeat run separate. The first includes account selection and connection overhead. The repeat run is the better representation of ordinary use.

One concrete case: a backyard movie night

Use one project across all three services: prepare a backyard movie night for eight adults. The project needs groceries, a simple invitation flyer, and approximately two hours of background music.

Fixed conditions

  • eight adult guests;

  • no alcohol and no real order placement;

  • healthy snacks and non-alcoholic drinks;

  • a retro-style flyer without personal photographs;

  • a two-hour mix of instrumental lo-fi and modern city pop;

  • no publication, public sharing, or sending to other people;

  • the test ends at the predefined verification point for each service.

This case exposes three different kinds of boundary. Instacart approaches an action that spends money. Canva produces a creative object that may need selection, saving, editing, downloading, or publishing. YouTube Music produces media that may need saving, opening, playback, and manual correction.

Do not change the guest count, dietary constraint, design brief, or playlist duration between the manual and connected-app runs. A changed task invalidates a direct comparison.

How to count steps, time, and confirmations

What counts as one manual action

One manual action is one intentional user operation that changes the state of the process. Count each of the following separately:

  • submitting the initial prompt;

  • submitting an answer to a clarification question;

  • choosing an app, account, store, design candidate, or playlist;

  • pressing a button such as Connect, Link, Continue, Open, Save, or Play;

  • approving an account permission;

  • opening the target application;

  • changing a required option or correcting the generated object;

  • pressing Checkout, Download, Share, Publish, or Place order, if such an action is deliberately included in a different test.

Do not count scrolling, reading, moving the pointer, waiting for generation, or correcting a typo before submitting. If one button opens a new page and completes a selection, it is still one action. If the new page then requires account selection and consent, those are additional actions.

How to handle clarifications

Every submitted clarification is a manual action. Record its subject as well as its count. “Which store?” and “May I substitute unavailable products?” reveal different missing information and may point to different opportunities for improving the request.

Do not silently rewrite the prompt during a measured run. Finish the run, record the clarification, and test an improved prompt later as a separate variant.

When to start and stop the timer

Use a normal stopwatch. Start it when you submit the first request. Stop only at the service-specific verification point:

  • Instacart: the cart is open in Instacart, and its products and quantities can be inspected;

  • Canva: the selected design exists in the Canva account and can be edited;

  • YouTube Music: the created object is open in YouTube Music and can be played.

Do not stop when AI Mode says that the task is complete. The target is an inspectable object in the external service, not a completion sentence in Search.

What counts as confirmation

An approval gate is a point where progress pauses until the user grants access, chooses an option, or authorizes an external action. Classify every confirmation:

          Confirmation type
          Example
          Why it is separate




          Access
          Connect a Google account to a third-party account
          Usually associated with the first run or renewed consent


          Choice
          Select a store, substitution policy, or design candidate
          The system lacks a preference or must not assume one


          Handoff
          Open Instacart, Canva, or YouTube Music
          Work continues outside Search


          Save or playback
          Save a design, save a playlist, or press Play
          The object may exist only provisionally before this action


          Consequential action
          Place an order, publish a design, or share content
          It affects money, data, visibility, or another person
Enter fullscreen mode Exit fullscreen mode

Result codes

Assign one primary code to each run:

  • S — the target object was created and passed verification;

  • P — only a preview or incomplete object was produced;

  • H — the process handed unfinished work to the user in the target app;

  • U — the connected-app function was unavailable;

  • E — a technical error prevented completion;

  • Q — one or more clarification requests materially changed the run.

You may record a secondary code when useful, such as P/Q for an incomplete result that also required clarification. Define that convention before the test and use it consistently.

Preparing the accounts and observation sheet

Use the same computer, browser, network, Google account, interface language, and test date for all three connected-app runs. Disable browser extensions only if you are prepared to disable them for every compared run.

Check availability before measuring performance

  • Sign in to the Google account intended for the test.

  • Open Google Search and enter AI Mode through the interface available to that account.

  • Confirm that the conversation interface accepts the language you plan to use.

  • Check which of the three services are already connected.

  • Record the browser, operating system, device type, account type, interface language, and date.

  • Open a separate blank observation sheet or prepare a paper counter.

Do not treat a missing app button as a failed performance run. Record U and the environment details. You cannot measure steps or elapsed completion time for a workflow that was never offered.

Account connection and privacy

Account linking may use OAuth or another delegated authorization flow. Read each permission screen before continuing. Confirm the domain, selected account, requested permissions, and application name.

Do not put passwords, payment-card details, access tokens, private documents, another person’s information, or a full delivery address into the prompt. If Instacart needs a delivery area to show inventory, provide it through the service’s own interface and exclude it from screenshots and notes.

If you record the screen, stop recording before entering credentials or opening account-security pages. A written tally of actions is enough for this experiment.

Observation record

Create one record for each run:

Service:
Run ID:
Run type: manual baseline / first connected run / repeat connected run
Date and local time:
Browser and device:
Google account type:
Target app account:
Prompt version:
Start time:
Stop time:
Elapsed time:
Manual actions:
Clarification count:
Access confirmations:
Choice confirmations:
Handoffs:
Save or playback confirmations:
Consequential actions:
Agent stopping point:
Result code:
Object created:
Verification outcome:
Differences from the request:
Errors or retries:
Notes:
Enter fullscreen mode Exit fullscreen mode

Give prompts version numbers such as GROCERY-V1, FLYER-V1, and MUSIC-V1. This prevents a later edit from being mistaken for the original test input.

Run a manual baseline first

A benchmark needs a comparison point. Here, the manual baseline is the same task completed directly inside the target service without AI Mode.

Without a baseline, “five actions” has no useful meaning. The direct workflow might require four actions or forty. Only a like-for-like comparison shows whether the connection removed work.

Baseline procedure

  • Open the target service directly.

  • Start the stopwatch at the first task-related action.

  • Complete the same task under the same fixed conditions.

  • Count actions using the definitions above.

  • Stop at exactly the same verification point used for the connected run.

  • Check the resulting object against the same content requirements.

  • Record corrections separately from ordinary navigation.

The baseline does not have to be elegant. It must be comparable. Do not compare a text-only grocery suggestion from Search with a fully inspected and purchased order, or an image preview with a finished editable design.

Steps saved: baseline manual actions − connected-run manual actions.

Time saved: baseline elapsed time − connected-run elapsed time.

Share of steps removed: steps saved ÷ baseline actions × 100%.

A negative result is valid. If the connected workflow requires more actions, enter the negative difference instead of converting it to zero.

Scenario 1: Instacart cart without placing an order

The target is an inspectable grocery cart for eight adults. The run ends before checkout. Do not assume that a text list in Search is a cart, and do not assume that a cart is ready merely because Instacart opened.

Reproducible request

Create an Instacart grocery cart for a backyard movie night
for 8 adults tonight. Include sparkling water, unsweetened drinks,
fresh fruit, vegetables with hummus, popcorn, and one nut-free snack.
Do not place the order. Ask before making substitutions.
Enter fullscreen mode Exit fullscreen mode

First connected run

  • Prepare the action counter and stopwatch.

  • Submit the request and start the timer.

  • If the interface asks you to select Instacart, count that selection.

  • If Link, Connect, or Continue appears, classify each required operation.

  • Read the authorization screen before approving access.

  • Record requests for a store, delivery area, delivery window, budget, or substitution policy.

  • Do not combine several submitted answers into one counted action.

  • Follow the handoff to Instacart if required.

  • Stop the timer only when the cart’s items and quantities can be inspected.

  • Do not proceed to checkout or place an order.

Cart verification checklist

  • All six requested categories are represented.

  • Quantities are plausible for eight adults.

  • No alcohol or unrelated product was added.

  • A nut-free snack is identifiable from the available product information.

  • No unapproved substitution was silently treated as accepted.

  • The selected store is visible.

  • Prices and fees are distinguished where the service displays them.

  • The cart remains available after one page refresh.

  • No order was placed.

Record the stopping point

Do not write merely “Instacart required confirmation.” Record the last completed operation and the next blocked operation:

Last completed operation: ______________________________
Next operation offered: ________________________________
Who had to perform it: AI Mode / user / Instacart
Reason for stopping: access / choice / handoff / checkout / error
Enter fullscreen mode Exit fullscreen mode

If AI Mode produces only a shopping list and the user must manually add every item, classify the run as a preview or handoff rather than a completed cart. If a cart exists but violates an important dietary constraint, use P and describe the defect.

Scenario 2: Canva flyer draft

The target is an editable flyer stored in the intended Canva account. A rendered candidate displayed inside Search is not enough unless you can open the same object in Canva and edit it.

Reproducible request

Create a flyer draft in Canva for a backyard movie night.
Use a retro 1980s cinema style, a dark navy background, warm yellow
headings, subtle film grain, and clear placeholders for date, time,
address, and movie title. Do not publish or share the design.
Enter fullscreen mode Exit fullscreen mode

First connected run

  • Submit the request and start the stopwatch.

  • Count a Canva selection if the service is not selected automatically.

  • Record every Link, Connect, account-selection, consent, or Continue action.

  • Record how many candidates appear, but do not count passive viewing as an action.

  • Inspect the candidates against the fixed brief.

  • Count the click used to choose one candidate.

  • Open the selected result in Canva.

  • Confirm that the design exists in the intended account.

  • Confirm that its text elements can be edited.

  • Refresh the page once and verify that the design remains accessible.

  • Stop the timer after those checks succeed.

  • Do not download, publish, share, or send the design.

Design verification checklist

  • There are separate placeholders for the date, time, address, and movie title.

  • The background and heading colors follow the request.

  • The retro cinema direction is recognizable without using personal material.

  • The primary text remains readable.

  • No invented real date, address, event name, or personal photograph appears.

  • The design can be edited without generating a new candidate.

  • The object survives a page refresh.

  • The design has not been published or shared.

Possible confirmation points

Record whether the user had to choose a candidate, explicitly save it, open Canva, select an account, or approve a format. These are different boundaries. A candidate displayed in Search but absent from the Canva account should receive P, not S.

If detailed text editing is possible only after the handoff, record the boundary as “editable object created; detailed editing left to user.” Do not count all later creative work unless it is necessary to meet the predefined completion criterion.

Scenario 3: YouTube Music two-hour playlist

The target is a named music object that opens in YouTube Music and can be played. Distinguish a saved playlist, an automatically generated mix, and a temporary list of recommendations. Record the object type exactly as the interface presents it.

Reproducible request

Create a 2-hour YouTube Music playlist for a backyard movie night.
Mix instrumental lo-fi beats with modern city pop. Keep the first
30 minutes calm, make the middle more upbeat, and avoid explicit tracks.
Save it as "Backyard Movie Night Test".
Enter fullscreen mode Exit fullscreen mode

First connected run

  • Submit the request and start the timer.

  • Count a YouTube Music selection if it is required.

  • Record account selection, connection, and consent actions separately.

  • Record whether the response creates a playlist, creates a mix, or returns recommendations.

  • Open the result in YouTube Music.

  • Check its displayed name and object type.

  • Check whether the result is available from the account after refreshing the page.

  • Inspect track count, total duration where available, genre mixture, ordering, and explicit labels.

  • Count Play as a separate confirmation if playback requires a user action.

  • Stop the timer when the verified object begins playback.

Playlist verification checklist

  • The displayed title matches “Backyard Movie Night Test.”

  • The total duration is close enough to two hours under a tolerance defined before testing.

  • Both instrumental lo-fi and modern city pop are represented.

  • The beginning is calmer than the middle according to a consistent manual review.

  • No visibly marked explicit track is present.

  • The same object can be opened again after a refresh.

  • The result is not merely a prose list of recommendations.

  • No content was shared with another person.

Define the duration tolerance first

“Two hours” rarely means exactly 02:00:00 at track level. Before running the test, define an acceptable range, for example:

Minimum acceptable duration: __:__:__
Maximum acceptable duration: __:__:__
Enter fullscreen mode Exit fullscreen mode

Choose your own tolerance before viewing the result. Do not widen it afterward to turn a failed result into a pass.

Record editing limits as observations

If you cannot add, remove, or reorder individual tracks through AI Mode, record where editing becomes manual. Do not assume that all playlist types or accounts behave identically. The useful observation is the exact operation attempted and what the interface offered next.

Repeat the run after the apps are connected

The repeat run separates normal operating cost from first-time authorization overhead. Keep each app connected, wait until the first result is safely recorded, and start a new AI Mode conversation.

  • Confirm that the same Google and target-service accounts are active.

  • Open a new conversation so the previous response does not supply hidden context.

  • Use the identical versioned prompt.

  • Start the timer at submission.

  • Count every manual action again from zero.

  • Stop at the same service-specific verification point.

  • Check whether the process created a duplicate object.

  • Record any renewed access request instead of omitting it as an anomaly.

If the complete linking flow appears again, the repeat run is still valid. Record it. The connection may not have persisted, consent may have expired, or the selected account may have changed.

For a more stable comparison, conduct three repeat runs per service and report the middle elapsed result after sorting the successful times. Do not mix unavailable or failed runs into a time calculation; show their result codes separately.

Comparison table: Instacart, Canva, and YouTube Music

Enter only observed values. A dash means “not measured,” not zero. Use mm:ss for elapsed time and integer counts for manual actions.

            Service
            Manual baseline
            First connected run
            Repeat connected run
            Confirmations observed
            Agent stopping point
            Code


            Steps
            Time
            Steps
            Time
            Steps
            Time




            Instacart
            ___
            __:__
            ___
            __:__
            ___
            __:__

              Access: ___
              Store or item choice: ___
              Handoff: ___
              Checkout: not performed

            Last automatic action: __________Next manual action: __________
            ___


            Canva
            ___
            __:__
            ___
            __:__
            ___
            __:__

              Access: ___
              Candidate choice: ___
              Save or handoff: ___
              Publication: not performed

            Last automatic action: __________Next manual action: __________
            ___


            YouTube Music
            ___
            __:__
            ___
            __:__
            ___
            __:__

              Access: ___
              Save: ___
              Handoff: ___
              Play: ___

            Last automatic action: __________Next manual action: __________
            ___
Enter fullscreen mode Exit fullscreen mode

Calculated savings

            Service
            Baseline steps
            Repeat-run steps
            Steps saved
            Baseline time
            Repeat-run time
            Time saved
            Object passed verification?




            Instacart
            ___
            ___
            ___ − ___ = ___
            __:__
            __:__
            __:__
            Yes / No / Partial


            Canva
            ___
            ___
            ___ − ___ = ___
            __:__
            __:__
            __:__
            Yes / No / Partial


            YouTube Music
            ___
            ___
            ___ − ___ = ___
            __:__
            __:__
            __:__
            Yes / No / Partial
Enter fullscreen mode Exit fullscreen mode

Confirmation map

          Service
          Preparation target
          Potential user decision to observe
          Consequential action excluded from this test




          Instacart
          An inspectable grocery cart
          Account, store, location, substitutions, or handoff
          Checkout and order placement


          Canva
          An editable flyer in the intended account
          Candidate selection, saving, format, or handoff
          Publication, download, or sharing


          YouTube Music
          A named, playable music object
          Account, save, handoff, playback, or manual track editing
          Sharing with another person
Enter fullscreen mode Exit fullscreen mode

This confirmation map is a test checklist, not a claim that every listed prompt will appear. Mark only the confirmations that your run actually required.

Verify that the task was actually completed

The AI Mode response is not proof of completion. Verification must happen in the system where the resulting object is supposed to exist.

Three levels of verification

  • Existence: the cart, design, or music object is visible in the target service.

  • Persistence: the same object remains available after a refresh or a new visit.

  • Conformance: the object satisfies the measurable requirements of the original request.

          Verification
          Instacart
          Canva
          YouTube Music
    
          Object exists
          The cart opens
          The design appears in the intended account
          The playlist or mix opens
    
          Object persists
          Items remain after refresh
          The selected design remains after refresh
          The same object can be opened again
    
          Content conforms
          Categories, quantities, and dietary constraint pass
          Style, placeholders, and editability pass
          Duration, genres, sequence, and explicit-content rule pass
    
          Excluded action did not occur
          No order was placed
          No design was published or shared
          No content was sent to another person
    

An existing object can still fail the test. A cart containing an unsuitable snack, a flyer without a date placeholder, or a playlist far outside the declared duration range is a partial result. Record P and the correction required.

Duplicate check

The repeat run may create another cart, design, or playlist. Record whether duplication occurred and whether the interface warned you. A faster repeat run is not necessarily better if every request leaves an unwanted duplicate that must be cleaned up manually.

Duplicate created: yes / no / unclear
Warning displayed before duplication: yes / no
Cleanup actions required: ___
Cleanup time: __:__
Enter fullscreen mode Exit fullscreen mode

Failure cases and what they mean

The app is not offered

Record U, the account type, language, device, browser, region, and date. Do not assign zero seconds or zero steps. The workflow was unavailable, not instantaneous.

Search returns instructions but no connected action

Confirm that the app name appears explicitly in the request. If the same versioned prompt still produces only prose, preserve the response description and assign U or P according to your predefined rule.

The wrong account opens

Stop before saving or creating private material in the wrong account. Record the run as an error, note where account identity became visible, and start a new run only after correcting the account selection.

Link and Continue repeat in a loop

Do not approve the same screen indefinitely. Record the number of cycles, assign E, and inspect the connection state. One controlled retry in a new conversation is enough for the primary test.

The result appears in Search but not in the app

Classify the result as P. For Canva, verify whether a specific candidate had to be selected. For YouTube Music, distinguish a saved playlist from a temporary mix or recommendations. For Instacart, distinguish an actual cart from a formatted shopping list.

The agent asks many questions

Count each submitted response. Do not simplify the prompt during the measured run. Later, create a second prompt version that supplies the missing information and test it independently.

The content is poor but the integration succeeded

Separate workflow completion from content quality. An object may have been created successfully while failing the brief. Record the stopping point and action count normally, then assign P for verification.

A handoff opens the service but loses the result

Record where the context was lost. Opening an empty Canva home page, an unrelated YouTube Music screen, or an Instacart page without the prepared items is not a completed handoff.

The repeat run reuses the old object

Determine whether the service reopened the first result or created a new one. A repeat test must not be marked successful merely because an earlier object remains available. Record the object name, visible identifier where safe, or creation time so the runs can be distinguished.

A consequential action appears ready to execute

Stop. Record the button label and the information shown immediately before it. Do not press Place order, Publish, Share, or a similar final action in this protocol.

Limitations of the test

  • Availability changes. Different accounts may receive different connected-app options.

  • Interface text changes. Button names and handoff sequences may not match another tester’s screen.

  • First runs cost more. Account selection and authorization should not be attributed to every future task.

  • Network conditions affect time. Use the same connection and report multiple repeat runs when possible.

  • Catalogs differ. Grocery inventory and music availability depend on the target service and account context.

  • Quality is not automation. Fewer actions do not guarantee a correct cart, usable design, or suitable playlist.

  • Handoff is not automatically a failure. It may represent a service boundary or a deliberate confirmation point.

  • One user’s time is not universal. Publish the environment and test date with any measured values.

  • Manual counting has error. Screen recording or a second observer can improve step classification, provided sensitive information is protected.

  • The three tasks are not equally complex. Compare each app with its own manual baseline, not directly with another app’s raw time.

  • No purchase or publication is tested. This protocol cannot prove what happens after the final excluded confirmation.

If a service fails or is unavailable, keep that run outside successful-time calculations. Report it as E or U. Averaging a failure into elapsed time hides the information readers need.

How to draw a conclusion without overstating the result

A connected app is useful in this experiment only if the repeat run reduces manual steps or elapsed time and the created object passes verification. An app card, generated preview, or confident completion message is not sufficient evidence.

Write one conclusion per service using the same structure:

In the repeat run, [service] changed the workflow from ___ to ___
manual actions and from __:__ to __:__. AI Mode stopped before
________________. The result [passed / partially passed / failed]
verification because ____________________________________________.
Enter fullscreen mode Exit fullscreen mode

Then add the confirmation evidence:

Access confirmations: ___
Choice confirmations: ___
Handoffs: ___
Save or playback confirmations: ___
Consequential confirmations tested: none
Unexpected manual corrections: ___
Enter fullscreen mode Exit fullscreen mode

The most valuable output is not a ranking of three brands. It is a handoff map. That map shows which preparation steps were removed, which decisions remained human, what existed in the target service, and which final action was intentionally left outside the agent’s control.

If the connected run saves time but produces an incorrect object, report both facts. If it saves no steps but reduces cognitive work, describe that benefit separately instead of converting it into an invented numeric result. If the function is unavailable, publish the availability result and environment without guessing what the workflow would have done.

Continue learning

Find more reproducible experiments in Agent Lab Journal Guides, and review the terms used in this test in the Agent Lab Journal Glossary.

Agent Lab Journal
Real experiments. Verifiable conclusions.
Enter fullscreen mode Exit fullscreen mode

Original article: https://agentlabjournal.online/en/google-search-connected-apps-test.html?utm_source=devto&utm_medium=referral&utm_campaign=agentlabjournal-en-global-all&utm_content=article&utm_term=google-search-connected-apps-test

Top comments (0)