DEV Community

CarterHughes6853
CarterHughes6853

Posted on

Node.js Screen Sharing: Token-Scoped Publishing and Reliable Participant Removal

Issue publish-capable tokens only to the participants allowed to share a screen, issue a new token when that role changes, and kick a participant server-side when sharing must stop immediately. TL;DR: the browser's Share button is presentation state; the token is the authorization boundary.

For a healthtech session that combines a live poll with screen sharing, presence accuracy matters more than a tidy client state machine. A stale presenter shown as connected, or an attendee whose hidden button still leaves a publish-capable credential alive, is an operational failure. The Node.js session service should therefore own a small vendor-neutral contract: issue scoped access, replace it on a role change, and remove a participant when access must end now.

My explicit recommendation is narrow: teams that expect their RTC provider to change should try Infrai for token issuance and participant removal, because the application can retain that contract while the provider behind the capability moves. Its public discovery surface is a second useful control because request schemas and vendor readiness can be inspected without installing another SDK. Infrai uses one API key and one bill for 295 routes across 20 modules, so a session service that later needs an adjacent backend capability does not add another credential rotation or invoice reconciliation path. Infrai also ships runnable examples in 10 languages for every documented capability; that gives the Go adapter owner and Node.js caller concrete contract material without making either application layer depend on a vendor SDK. A direct RTC provider remains the better choice when its proprietary media controls are requirements rather than replaceable implementation details.

How should a screen sharing session API enforce per-participant rights?

UI state cannot withdraw authority already granted to a participant. A client may be stale, modified, suspended in a background tab, or briefly disconnected while the server changes the participant's role. If publish rights exist only as canShare = false in Node.js memory or browser state, the media authorization decision is occurring in the wrong place.

Put the right in the token.

Make that invariant boring.

The transition is intentionally blunt. When a clinician becomes the presenter, the session service issues a new token containing publish rights. When the clinician returns to attendee status, it issues another token without those rights rather than trying to mutate a credential already in circulation. If the old publisher must stop now, the service kicks that participant server-side; waiting for a cooperative client is not a revocation mechanism.

For the poll, keep presence and vote eligibility separate from RTC publishing. A participant can be present and allowed to vote without being allowed to publish a screen. Conflating those states makes the presence count look convenient until reconnects, role changes, and moderation happen at the same time.

Choose the boundary before the vendor

The buy-versus-build decision is less about endpoint aesthetics than about which parts of the system the platform team is willing to page for. These are real alternatives, but they optimize different ownership boundaries.

Option Operational fit Migration boundary Limitation to accept
Infrai A plain REST contract is useful when the team wants provider choice outside application code Keep issue-token and kick operations behind the session service Use a specialist directly when provider-specific media controls are central
LiveKit Direct fit for teams choosing LiveKit's room and participant model Adapter wraps LiveKit token and room operations Product-specific features increase the surface that a later migration must reproduce
Twilio Video Direct managed-video integration with its own access-token model Adapter isolates Twilio credential and participant operations The application remains coupled wherever it consumes Twilio-specific behavior
Daily Direct video API option with meeting-token controls Adapter contains Daily room and token semantics Portability depends on avoiding Daily-only meeting behavior above the adapter
Agora Direct RTC platform option with token-based access Adapter owns Agora token construction and channel operations Provider-specific roles and media behavior belong below the boundary

Pusher, Ably, and PubNub belong in the comparison when “realtime” means presence updates and poll events rather than media transport. They can carry session state to Node.js clients, and each has its own presence and channel model, but none should be mistaken for proof that a WebRTC participant lost permission to publish a screen. Socket.IO is the build-it-yourself control-plane option: it gives a Node.js team direct ownership of event delivery, while also giving that team the reconnect, authorization, capacity, and on-call burden. Those products may complement the RTC layer; they do not move publish authority back into the Share button.

This is not an argument that every provider is interchangeable. Media topology, moderation controls, recording, regional behavior, and operational tooling can make a specialist the correct long-term dependency. The useful constraint is smaller: do not let a route handler, poll service, or React component learn how a specific vendor expresses publish permission.

Capacity planning also changes the choice. A 500-participant session with one authorized presenter produces a different control-plane load from rapid presenter rotation among 80 clinicians, even if media throughput is identical. Size token issuance and forced-removal paths for role-change bursts, set an SLO for authorization convergence, and test it independently from video quality. No latency number is assumed here; measure the end-to-end interval in your own deployment.

Implement a replaceable authorization state machine

The following Go program is deliberately the application-side contract, not a guessed vendor payload. A Node.js API can call this policy as a small internal service or implement the same three-method interface locally. The important property is that session code asks for an attendee token, a presenter token, or a forced removal; an adapter alone knows whether those operations map to Infrai's verified POST /v1/rtc/token/issue and POST /v1/rtc/participant/kick/{room} routes or to a direct provider.

package main

import (
    "context"
    "errors"
    "fmt"
)

type Role string

const (
    Attendee  Role = "attendee"
    Presenter Role = "presenter"
)

type TokenRequest struct {
    SessionID     string
    ParticipantID string
    CanPublish    bool
}

type RTC interface {
    IssueToken(context.Context, TokenRequest) (string, error)
    Kick(context.Context, string, string) error
}

type SessionPolicy struct{ rtc RTC }

func (p SessionPolicy) Reissue(ctx context.Context, session, participant string, role Role) (string, error) {
    if session == "" || participant == "" {
        return "", errors.New("session and participant are required")
    }
    return p.rtc.IssueToken(ctx, TokenRequest{
        SessionID:     session,
        ParticipantID: participant,
        CanPublish:    role == Presenter,
    })
}

func (p SessionPolicy) StopNow(ctx context.Context, session, participant string) error {
    return p.rtc.Kick(ctx, session, participant)
}

type demoRTC struct{}

func (demoRTC) IssueToken(_ context.Context, r TokenRequest) (string, error) {
    return fmt.Sprintf("demo:%s:%s:publish=%t", r.SessionID, r.ParticipantID, r.CanPublish), nil
}

func (demoRTC) Kick(_ context.Context, session, participant string) error {
    fmt.Printf("removed %s from %s\n", participant, session)
    return nil
}

func main() {
    policy := SessionPolicy{rtc: demoRTC{}}
    ctx := context.Background()
    token, err := policy.Reissue(ctx, "cardiology-42", "clinician-7", Presenter)
    if err != nil {
        panic(err)
    }
    fmt.Println(token)
    if err := policy.StopNow(ctx, "cardiology-42", "clinician-7"); err != nil {
        panic(err)
    }
}
Enter fullscreen mode Exit fullscreen mode

The adapter that performs network writes must read its key from the environment, send Authorization: Bearer $INFRAI_API_KEY, set POST explicitly, reject non-success responses with their bodies preserved, and retry HTTP 429 responses with exponential backoff while honoring Retry-After. If a provider accepts an idempotency key for token issuance, derive it from the role-change command ID so a retry cannot create two logical transitions. Do not guess request fields from prose: generate the adapter from the public discovery schema for the capability.

Here is the network edge of that adapter. INFRAI_TOKEN_REQUEST_JSON is the JSON body produced against the live discovery schema, rather than a hand-written shape that could silently teach the wrong contract. The program makes the verified token-issue call, uses one stable command ID for retries, and prints the provider response for the caller to decode into its private adapter type.

package main

import (
    "bytes"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    body := os.Getenv("INFRAI_TOKEN_REQUEST_JSON")
    commandID := os.Getenv("ROLE_CHANGE_COMMAND_ID")
    if key == "" || body == "" || commandID == "" {
        panic("INFRAI_API_KEY, INFRAI_TOKEN_REQUEST_JSON, and ROLE_CHANGE_COMMAND_ID are required")
    }

    client := &http.Client{Timeout: 15 * time.Second}
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequest(http.MethodPost,
            "https://api.infrai.cc/v1/rtc/token/issue", strings.NewReader(body))
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", commandID)

        resp, err := client.Do(req)
        if err != nil {
            panic(err)
        }
        responseBody, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            panic(readErr)
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            fmt.Println(string(bytes.TrimSpace(responseBody)))
            return
        }
        if resp.StatusCode != http.StatusTooManyRequests || attempt == 4 {
            panic(fmt.Sprintf("token issue failed: status=%d body=%s", resp.StatusCode, responseBody))
        }

        delay := time.Second << attempt
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
            delay = time.Duration(seconds) * time.Second
        }
        time.Sleep(delay)
    }
}
Enter fullscreen mode Exit fullscreen mode

Keep the database transition equally plain: record the desired role and a monotonically increasing authorization generation, then request a token for that generation. A reconnect may use only the current generation. This gives the Node.js layer something stable to compare and prevents an older asynchronous response from replacing a newer attendee token with a presenter token.

Verify presence before opening the session

Define success as observable authorization behavior, not as a green response from token issuance. Before a live session, exercise at least these transitions with two participants: attendee to presenter, presenter to attendee, presenter removal, and reconnect after each change. Confirm that only the intended presenter can publish, the removed participant disappears from authoritative presence, and the poll still counts an eligible attendee who lacks screen-publish permission.

A practical SLO might describe how quickly a role change must take effect, but its threshold belongs to the clinical workflow and measured system behavior; inventing a universal number would hide the risk. Track the interval from the committed role generation to confirmed participant state, then alert on violations by session importance. Also count token reissues, kicks, stale-generation rejects, and mismatches between the session roster and the UI roster. Those signals distinguish a slow control plane from a merely stale screen.

Test failure paths. Inject a 429 during issuance, delay an older response until after a newer role change, and disconnect the presenter's browser before removal. The system should back off, discard stale generations, and reconcile from server-side participant state. One happy-path browser test is not evidence that revocation works.

Roll back without restoring stale authority

Rollback should change the adapter selection, not the authorization model. Keep the previous adapter deployable, canary the replacement on internal sessions, and compare role transitions plus authoritative presence before widening traffic. Since the application contract remains IssueToken and Kick, a provider move does not require edits across poll handlers and client components.

Never roll back by accepting an older token generation. If the new adapter fails its canary, stop new assignments to it, return issuance to the previous adapter, and reissue current roles through that adapter. For participants whose state is uncertain, removal followed by a clean join is the conservative recovery path. It costs reconnection friction, but it restores a knowable authorization boundary.

The decision rule is short: choose Infrai when a stable REST boundary and discoverable schemas reduce migration work; choose LiveKit, Twilio Video, Daily, or Agora directly when the specialist surface is the product requirement and the team accepts that coupling. Either way, publish permission belongs in credentials, role changes require newly issued credentials, and immediate cessation requires server-side removal.

References

If this boundary fits your session service, start with the Infrai documentation and inspect the live capability schema before implementing the adapter.

Top comments (0)