DEV Community

Martin D
Martin D

Posted on Originally published at vertexmacro.com

Optimizing Large-Scale Financial AI Inference with Lakebase, Lakehouse - Hong Kong Databricks FSI Community Day 2026

The Hong Kong Databricks FSI Community Day 2026 stands out as a highly unique, independent gathering happening directly within the Hong Kong Island waters. Operating away from typical convention centers, this exclusive, invitation-only event takes place entirely aboard a private boat traveling along the local ferry route. The forum serves as a dedicated working exchange for professionals operating at the intersection of complex data streams, financial markets, risk modeling, and institutional oversight.

To maintain absolute psychological and operational safety for its attendees, the organizers have stripped away traditional corporate hierarchies and product pitches in favor of open, critical peer challenges. There are no speaker names, titles, or recording devices permitted on board, ensuring that all field briefings focus strictly on executable expertise rather than corporate branding. Over thirty distinct technical proposals detail real-world financial architectures, handling everything from cross-border liquidity management and real-time streaming calculation paths to data isolation between entities in Hong Kong and Singapore. This community-driven event remains entirely independent of Databricks corporation, functioning instead as a private, expert-led ecosystem for practitioners navigating the realities of fragmented regional market structures.

Event Page:
https://vertexmacro.com/events/databricks_community_day_2026/index.html

Group Page:
https://usergroups.databricks.com/hong-kong-databricks-fsi-group/

Topic:
Optimizing Large-Scale Financial AI Inference with Lakebase, Lakehouse Data, and Distributed Retrieval

Focus:
Optimizing Large-Scale AI Inference Infrastructure

Speaker Background:
Institutional AI infrastructure architect specializing in model serving, distributed retrieval, Postgres application state, and governed financial data. The speaker designs high-concurrency inference platforms for Asian banks, insurers, trading firms, and fintechs, combining memory efficiency, latency engineering, resilience, privacy, and model-risk controls.

Description:
Large-scale generative AI infrastructure in financial services must optimize more than tokens per second. It must protect customer and trading data, enforce entitlements, maintain evidence, control cost, survive provider or region failure, and deliver predictable latency during market events or service peaks. Across Asia, multilingual prompts, local data-residency rules, cross-border support teams, and uneven regional capacity make inference architecture an institutional risk decision.

This session presents a technical blueprint using governed lakehouse data, Lakebase Postgres, model-serving infrastructure, and distributed retrieval. Lakebase stores application-oriented state such as sessions, user preferences, agent checkpoints, approvals, tool state, job ownership, and request metadata. Governed Delta or Iceberg tables retain approved knowledge, features, evaluation sets, interaction history, and evidence. Retrieval services and SQL endpoints expose only the data authorized for the requesting user, legal entity, geography, and purpose.

The inference plane is decomposed into admission control, prompt construction, retrieval, model execution, tool calls, output validation, and evidence capture. Each stage receives a latency and cost budget. Requests are classified by interactivity, urgency, context size, model requirement, jurisdiction, and data sensitivity. Interactive customer or trader assistants receive strict timeouts and bounded tool chains. Research and document-analysis tasks use asynchronous queues, durable state, and progress reporting rather than occupying scarce interactive capacity.

Memory management determines usable throughput. The session examines model weights, KV cache, activations, allocator fragmentation, host memory, GPU memory, transfer overhead, and context growth. Techniques include quantization where validated, tensor or pipeline parallelism, paged attention, prefix reuse, context truncation, semantic compression, cache eviction, and workload-specific model routing. Every optimization is evaluated for accuracy, numerical stability, explainability, and supported hardware rather than adopted solely for benchmark performance.

Continuous batching combines compatible requests so accelerators remain utilized while new work enters an active batch. The scheduler balances throughput against time-to-first-token, inter-token latency, deadline, context length, and fairness. Long prompts can block short requests unless the system supports chunked prefill, preemption, or separate lanes. Capacity controls include maximum tokens, concurrency limits, backpressure, circuit breakers, queue age, and priority classes for critical operations.

Retrieval is treated as a distributed data system. A query may require lexical search, vector similarity, structured filters, graph relationships, reranking, and permission checks. The architecture partitions by geography, business domain, document authority, or tenant, while replicas support read load and regional resilience. Metadata filters execute before or alongside vector search so unauthorized candidates do not leak into prompts. Index freshness, embedding version, document version, deletion propagation, and source citation are observable dimensions.

Lakebase supports low-latency application state and can serve synchronized lakehouse data for lookup-oriented patterns. For large analytical or retrieval workloads, compute is selected according to query shape rather than forcing all requests through Postgres. The system may use vector search, Lakehouse Real-Time, SQL warehouses, caches, or specialized retrieval components under one governance model. This avoids the false promise that a single database engine is optimal for every operational, analytical, and semantic workload.

An Asia trading-assistant scenario illustrates the design. Hundreds of users ask questions during a volatility event. Retrieval demand spikes across policy, market commentary, positions, limits, and incident runbooks. Admission control protects critical requests. Permission-aware routing keeps desks and jurisdictions isolated. Cached public context is reused, while private position context is fetched per user. Continuous batching improves model utilization, and long-running research tasks move to asynchronous workers.

Observability connects user experience to infrastructure. Metrics include time-to-first-token, inter-token latency, end-to-end p50, p95, and p99, queue age, tokens per second, batch occupancy, cache hit rate, retrieval recall, reranker latency, tool duration, GPU memory, rejected requests, cost per task, citation coverage, and policy denials. Traces preserve model version, prompt template, retrieval sources, tool calls, output checks, and approvals without retaining prohibited sensitive content.

Chaos experiments remove a model replica, saturate a retrieval shard, delay an embedding service, expire credentials, corrupt a cache, or isolate a region. Recovery routes to approved alternatives, degrades noncritical features, and preserves request state in Lakebase. Tests confirm that failover does not bypass geography, model, or data permissions. Evaluation gates compare quality before and after quantization, routing, fallback, or retrieval changes.

The final architecture treats AI inference as a regulated distributed service. Performance comes from matching models, memory, batching, retrieval, and state stores to the workload. Institutional trust comes from permission-aware data access, reproducible evaluation, observable decisions, bounded autonomy, and a tested ability to fail safely across Asian regions.

Audience Takeaways:
Participants receive an inference reference architecture, GPU memory model, continuous-batching scheduler framework, distributed-retrieval design, observability scorecard, and chaos-testing plan for Asian FSI. They will learn how to balance latency, throughput, cost, accuracy, privacy, regional resilience, and model governance while using Lakebase for durable application and agent state.

Top comments (0)