The Hong Kong Databricks FSI Community Day 2026 stands out as a highly unique, independent gathering happening directly within the Hong Kong Island waters. Operating away from typical convention centers, this exclusive, invitation-only event takes place entirely aboard a private boat traveling along the local ferry route. The forum serves as a dedicated working exchange for professionals operating at the intersection of complex data streams, financial markets, risk modeling, and institutional oversight.
To maintain absolute psychological and operational safety for its attendees, the organizers have stripped away traditional corporate hierarchies and product pitches in favor of open, critical peer challenges. There are no speaker names, titles, or recording devices permitted on board, ensuring that all field briefings focus strictly on executable expertise rather than corporate branding. Over thirty distinct technical proposals detail real-world financial architectures, handling everything from cross-border liquidity management and real-time streaming calculation paths to data isolation between entities in Hong Kong and Singapore. This community-driven event remains entirely independent of Databricks corporation, functioning instead as a private, expert-led ecosystem for practitioners navigating the realities of fragmented regional market structures.
Event Page:
https://vertexmacro.com/events/databricks_community_day_2026/index.html
Group Page:
https://usergroups.databricks.com/hong-kong-databricks-fsi-group/
Topic:
Scaling Permission-Aware FSI Agent Inference Across Hong Kong and Singapore
Focus:
Optimizing Large-Scale AI Inference Infrastructure
Speaker Background:
Institutional AI infrastructure architect specializing in high-concurrency inference, distributed retrieval, identity-aware agents, and governed financial data. The speaker designs secure platforms for Asian banks, insurers, and trading firms, combining GPU efficiency, low-latency retrieval, legal-entity isolation, resilience, privacy, cost control, and model-risk governance.
Description:
Large-scale financial AI inference is not safe merely because the model endpoint is private. A request may cross retrieval indexes, SQL engines, caches, prompt stores, tool servers, model providers, traces, and regional failover systems. If any component loses the caller’s legal-entity context, an HK agent may retrieve SG customer information or infer another branch’s activity from cache keys, counts, timing, or errors.
This session presents a permission-aware inference architecture for Agent Bricks and Genie-based experiences across Entity HK and Entity SG. Each deployed agent receives a dedicated service principal, bounded tools, approved models, entity groups, regional configuration, and cost budget. Human users retain their own identity where end-user authorization is required. Agent identity never replaces user identity silently; the effective authorization model defines whether privileges are intersected, delegated, or service-owned.
Unity Catalog provides the governed data foundation. The three-level namespace organizes assets, while object privileges establish base access. Governed tags mark jurisdiction, entity, residency, PII, confidentiality, purpose, and authoritative status. ABAC policies dynamically apply row filters and column masks. A Hong Kong retrieval query therefore receives only permitted rows, while customer name, account, counterparty identifiers, and commercial secrets are masked according to role and purpose.
Policy enforcement is described precisely. Unity Catalog evaluates policy scope, principals, group membership, and tags. Databricks Runtime enforces the effective filter or mask during query execution. The architecture does not rely on an undocumented “entitlement interceptor,” nor does it claim that Photon file pruning is the legal control. Query optimization can reduce reads, but the security guarantee comes from governed privileges and runtime-enforced policies.
Genie Ontology gives agents a business-aware map of certified metrics, domains, authoritative sources, and business rules. HK and SG can share group-level terminology while retaining entity-specific definitions. A term such as “available liquidity” may use different source systems, cut-off times, currencies, or regulatory deductions. Permission-aware ontology context prevents the agent from resolving an HK question using an SG source that the caller cannot access.
The inference path contains admission control, identity resolution, policy context, prompt construction, retrieval, reranking, model execution, tool calls, output validation, and evidence capture. Every stage carries entity and jurisdiction context. Request classification separates interactive assistance, long-running research, customer support, risk analysis, and regulated decision support. Bounded tool chains, token limits, timeouts, circuit breakers, and human approval prevent an agent from expanding its scope through repeated planning.
GPU optimization remains important. The session covers weight placement, KV cache, context growth, allocator fragmentation, quantization, paged attention, prefix reuse, tensor parallelism, model routing, and continuous batching. Batches are grouped only when isolation guarantees are preserved. Cache keys include authority context, model, prompt policy, data version, and entity. Private retrieved context is never reused across legal entities merely because prompts are semantically similar.
Distributed retrieval uses entity-aware partitions, replicas, and regional placement. Metadata filters and permissions are evaluated before or alongside vector search so unauthorized candidates never become model context. Index freshness, embedding version, deletion propagation, document authority, citation, and residency are observable. Shared public content can be cached globally; private position, client, and transaction context remains entity-scoped.
A volatility event generates hundreds of concurrent requests from Hong Kong and Singapore trading and risk teams. Admission control reserves capacity for critical risk queries. Continuous batching improves accelerator utilization, while fairness prevents one entity from exhausting capacity. Long-running research moves to asynchronous workers with durable state. If an SG retrieval shard fails, routing selects only approved SG alternatives; it never falls back to HK data to preserve availability.
For group-level macro statistics, Clean Rooms can provide an isolated no-trust collaboration environment where approved notebooks operate over shared assets without direct raw-data access. Aggregate-only code, output review, threshold controls, and optional differential privacy must be designed and approved. The platform should not claim that every clean-room query automatically receives differential-privacy noise.
Observability records principal, user, entity, model, retrieval sources, policy decisions, mask application, cache scope, tool calls, tokens, latency, cost, citations, denials, and approval. Traces avoid prohibited prompt content while preserving reproducibility. Red-team and chaos exercises attempt cross-entity prompt injection, cache poisoning, confused-deputy tool use, stale entitlement, region loss, and model fallback.
The final architecture optimizes throughput only inside verified legal boundaries. Success means predictable latency and high utilization without cross-entity retrieval, inference, caching, or failover. Institutional-grade AI requires every answer to be attributable to an authorized identity, permitted source, approved model, documented policy, and reviewable decision path.
Audience Takeaways:
Participants receive a permission-aware inference architecture, agent identity model, Unity Catalog policy pattern, Genie Ontology design, entity-safe batching and caching strategy, distributed-retrieval blueprint, and red-team plan. They will learn how to scale HK-SG AI workloads while preserving residency, confidentiality, human approval, model governance, observability, and safe regional failover.
Top comments (0)