Release week. Apache Iceberg 1.12.0 passed its vote after a license audit sank the first candidate, Apache Polaris shipped 1.8.0 with a security fix inside it, and Iceberg C++ 0.4.0 went out the door. Under the release traffic sat a quieter theme that ran through every list. Specs are getting stricter about what they promise. Iceberg voted to forbid new equality deletes in v4. Parquet argued over whether a vector type should reject NaN. Ossie asked what a column name in a metric expression points to. Polaris debated whether a config field should be an enum or free text. Each thread is a fight about contracts, and each one ends with a narrower, more checkable promise to the engines that read these formats.
This issue covers roughly 380 messages across the Iceberg, Polaris, Arrow, Parquet, DataFusion and Ossie dev lists between September 23 and September 30.
Apache Iceberg
Iceberg 1.12.0 passes on the second try
Neelesh Salian ran the 1.12.0 release, and it took two candidates to land. The RC1 vote drew early +1s from Alex Dutra, Gianluca Graziadei and Szehon Ho. Then Kurtis Wright posted a -1 with a disclaimer at the top: he had prompted Claude to audit every license in the candidate. The audit found that the GCP bundle jar shipped Caffeine with no entry in its LICENSE file. Kevin Liu followed up with his own table of gaps and found a second one around Jackson's FastDoubleParser. Jean-Baptiste Onofré cast a -1 on the same grounds. Neelesh sorted each report, landed the fixes, and pulled in Gianluca's patch for a variant-column NoSuchMethodError that hit clusters with older commons-lang3 jars.
The RC2 vote opened on September 25. Neelesh stretched the window to Tuesday afternoon to cover the weekend. It had one scare. Huaxin Gao pointed out that the staging repository had no Spark 4.2 artifacts, even though the source tree, CI matrix and Gradle defaults all pointed at 4.2. Amogh Jahagirdar agreed a new RC looked necessary. Russell Spitzer then asked why an earlier PR had excluded 4.2 from convenience binaries on purpose. Amogh found the older thread. The community had held back 4.2 binaries while Spark's own APIs kept moving. Huaxin withdrew the alarm and voted +1 with a note that every RC1 license gap was closed in the jars themselves.
The result came in on September 29: four binding +1s from Russell Spitzer, Kevin Liu, Szehon Ho and Amogh Jahagirdar, and ten non-binding +1s. Two details stood out in the votes. Kevin ran an agent over the repository looking for correctness bugs and flagged a handful of issues that also exist in 1.11. He called them non-blockers and listed them for triage. Kurtis came back with a second AI-assisted pass, this time labeled as the output of an agent skill built for checking open source release candidates. Release verification is turning into a place where contributors point agents at tedious, rule-driven work and then sign their names to the result.
The CI side of the release had its own story. Steven Wu asked where the MinIO replacement stood after Kafka Connect integration tests failed on the MinIO container. Neelesh opened a tracking issue and a stopgap PR. Within a few hours Kevin reported that the switch was done across every Iceberg repository, crediting Sreesh for the work. Amogh later noted that he ran the RC2 integration tests against RustFS. The S3 test backend for the project is now a Rust implementation.
v4 forbids new equality deletes
Huaxin Gao called a spec vote on PR #17783. The direction was already settled in July: v4 writers must not produce new equality deletes, readers still apply existing ones from v2 and v3 tables, and an upgrade stays metadata-only. This vote covered the exact spec wording. Russell Spitzer questioned whether a second vote was needed at all, then voted +1 anyway. Ryan Blue said the point was to affirm the decision and merge it. Anoop Johnson said the change simplifies the v4 design a lot. The vote passed with four binding +1s from Russell, Yufei Gu, Steven Wu and Ryan, plus five non-binding. Huaxin will merge the PR and follow up with a separate change that restructures how the spec presents deletion vectors next to the legacy delete formats.
This is the most consequential spec change of the month for engine builders. Equality deletes let streaming writers such as Flink record "delete the row where id = 7" without reading the table first. They also push work onto every reader, which has to join the delete file against data files at scan time. v4 closes that door for new writes. Streaming pipelines that rely on equality deletes need a plan for deletion vectors before they adopt v4.
Who controls access: the delegation header debate
A thread that started in Polaris landed on the Iceberg list. youngrae kim asked what a REST server should do when a client sends X-Iceberg-Access-Delegation but the server can satisfy none of the requested mechanisms. Picture vended credentials against an S3-compatible store with no STS, or remote signing on a server that does not implement it. Polaris returns HTTP 400. The spec says the server can supply access through any or none of the requested mechanisms.
Alex Dutra proposed spec wording that favors failing fast, and Yufei Gu agreed. Daniel Weeks, who wrote much of the REST spec, disagreed with the framing. His argument was that the catalog should be the sole authority on access, and security around physical data should not be a negotiation between client and server. The loose wording was deliberate. It lets a catalog return credentials or remote signing config, or return metadata only, without the spec forcing its hand. Alex countered that the header cannot express a common client capability, which is using the client's own storage credentials and skipping delegation entirely.
youngrae closed the loop on September 29. Polaris will keep its 400 as a documented implementation choice. The remaining ask is narrow: one sentence in the spec that says whether the header is a capability statement or a requirement, and what an absent header means. That clarification leaves the catalog's decisions untouched. It only tells clients and servers how to read the signal.
A related effort moved forward. Sung Yun and Prashant Singh had been discussing access delegation for the FILE type, where a governed table references files that engines need to read directly. Sung reported that the two of them converged their proposals onto PR #18080, which now carries the optional request and response parameters from both drafts. Steven Wu also announced that the Thursday biweekly sync will cover both the File and Vector types starting October 8. He said the File type discussion is close to consensus, so more of the time will go to the Vector type design doc from Yan Yan.
Labels, SQL, and the agent-shaped consumer
Labels landed in the REST spec recently. A catalog can now attach object-level and field-level labels to a load-table response. Andrei Tserakhau asked whether Iceberg should expose them as a .labels metadata table. Péter Váry had raised the underlying issue in review. Every metadata table today derives from table metadata. A .labels table carries catalog-owned data that varies by catalog and can be empty.
Ryan Blue questioned the premise that labels need a SQL surface at all. He saw them as context for engines, such as cost attribution logs or attaching an engine policy to a table. Andrei answered that some consumers only speak SQL. He named governance tooling and, specifically, LLM agents that explore a warehouse by writing queries. Péter argued that the proposed uses turn labels into full table metadata, which contradicts the read-path design that treats them as a current view with no persistence guarantees. Ryan then took apart Andrei's cost-attribution example. The sample query joined .labels to .files with no ON clause, which produced a cartesian join that only looked right because the filter selected one label.
Andrei conceded the join argument and pointed to an alternative. PR #18049 surfaces object and field labels through DESCRIBE, which gives SQL-only consumers, agents included, a way to see labels without a new metadata table. The metadata table is parked until someone shows a use case that needs joins. The episode is worth noting. "An AI agent needs to read this" showed up as a design argument on the Iceberg list, and the community answered it with the smallest surface that works.
Encryption keys move closer to the files they protect
Gábor Kaszab reopened the design of encryption key storage. Today table metadata carries an encryption-keys list that holds both data encryption keys for manifest lists and the key encryption keys that wrap them. Gábor wants to extend encryption to statistics files and v4 root manifests. His proposal stores each file's wrapped key directly in the metadata entry for that file, one key per file. KEKs stay in the shared list and are still referenced by key-id. The spec draft for statistics files lives in PR #17533.
Gidon Gershinsky agreed the DEK metadata belongs in per-file structures in v4. He added that reusing a DEK across files is possible but requires careful handling to avoid breaking AES-GCM, so the practical rule is a fresh random DEK per file. Xander Bailey wanted the one-key-per-file rule stated explicitly and asked for a common wrapped-key structure across file types, so implementations share code. Gábor's closing argument was about cleanup. With keys in a shared list, every new encrypted file type needs its own cleanup trigger when the file is dereferenced. With keys stored inline, the key disappears with the file. Xander found the argument convincing and asked for a Java draft. Gábor wants the design settled first. Xander also proposed a recurring encryption community sync to cover client library work, the KMS vending spec and v4 changes together.
Smaller spec decisions
rahul mahadev called a vote on a standard User-Agent format for REST clients. The format lists products from most specific to least, such as Spark/4.0.0 iceberg-spark/1.9.0 iceberg-java/1.9.0 (scala 2.13.16). The Iceberg library token must appear when the header is sent. Servers must not reject requests based on it. Xuanwo liked that the header stays separate from authentication and capability negotiation. Alex Dutra left more PR comments to resolve before the vote proceeds.
Eduard Tudenhöfner announced a rename in the new content stats: avg_value_size_in_bytes becomes total_bytes in PR #18308. Averages do not aggregate cleanly across files. Totals do. Daniel Weeks confirmed that the average is still recoverable from total_bytes and value_count, and noted that a total gives planners an explicit upper bound for size estimates. Ryan Blue gave a +1.
Shawn Chang asked why the index spec requires region files to be globally ordered and non-overlapping, given the write amplification and maintenance cost. Péter Váry explained the target. Once index metadata is cached, a single-key lookup should cost one small file read. Overlapping regions force readers to consult several files and reconcile results, and on object storage every extra file read adds latency.
Releases beyond Java
Manu Zhang announced Iceberg C++ 0.4.0 on September 26. The vote passed with binding support from Gang Wu, Renjie Liu and Kevin Liu. Renjie's verification ran 4,152 individual tests across 19 test executables and exercised REST and SQL catalog support, including a SQLite catalog run. Xuanwo suggested a LICENSE path fix for the MurmurHash3 files, and Manu opened an issue for the next release.
Matt Topol opened the vote for Iceberg Go v0.7.0 on September 28. Neelesh Salian, Andrei Tserakhau and Xuanwo voted +1. Xuanwo ran 3,344 tests on Go 1.25.9 and saw one Puffin checksum assertion fail on Go 1.27.1. The code still rejected the corrupted input, so he called it non-blocking.
Community and housekeeping
Robin Moffatt opened PR #18019 to finish the Kafka Connect LICENSE and NOTICE cleanup that Ryan Blue started in April. The PR regenerates runtime-deps.txt, aligns the connector's cloud SDK dependencies with the AWS, GCP and Azure bundles, and drops the Hive distribution. The timing matters. A Confluent Hub security scan flagged the Iceberg Kafka Connect listing, and Confluent's policy escalates flagged connectors toward removal unless the partner responds.
Gaspard Merten shared a PyIceberg PR that moves upsert comparisons out of Python loops and into PyArrow compute. He built it while working on a data lake for a financial organization, and he said the change cuts memory and CPU for wide-table upserts. He also noted that he wrote part of it with Claude and reviewed it himself several times.
Viktor Kessler announced two European community meetups in back-to-back days: Berlin on November 18 and Amsterdam on November 19. Both have open calls for talks.
Apache Polaris
Polaris 1.8.0 ships with a CVE fix
Jean-Baptiste Onofré's first 1.8.0 candidate did not survive. Alex Dutra voted -1 on rc1 because a test file added in PR #4823 carried an incorrect license header. JB cancelled the vote and took another pass at LICENSE and NOTICE. The rc2 vote passed with binding +1s from Alex Dutra, Yufei Gu, Dmitri Bourlatchkov and JB, plus non-binding support from Prithvi S, Ajantha Bhat and Ayush Saxena. Yufei's verification checked 442 artifacts and ran authenticated CLI smoke tests against built and staged servers.
JB announced 1.8.0 on September 28. The notable changes:
- Semantic model privileges. Semantic models now get dedicated privileges for list, create, read, update and drop. Admins grant them to catalog roles on a single model, a namespace or a whole catalog. This is Polaris preparing to govern Ossie-style semantic models the same way it governs tables.
-
CLI toggles for STS and KMS.
catalogs updateaccepts--no-stsand--no-kms, so operators can change those settings on an existing S3 catalog instead of recreating it. -
GCP federation auth. The CLI adds
gcpas an external catalog auth type for Iceberg REST federation. That lets operators create GCP-authenticated catalogs such as BigLake without passing Google credential secrets as command-line flags. -
Paging. A global
--page-sizeoption paginates list calls, when the server enables theLIST_PAGINATION_ENABLEDflag. -
JDBC schema selection. The relational JDBC backend reads its schema from the driver's
currentSchemaproperty, so the persistence layer no longer hardcodes a schema name.
A day later JB published CVE-2026-97395, rated important, for all Polaris versions before 1.8.0. A principal with permission to set table properties was able to put FileIO settings such as s3.endpoint into table metadata. During server-side operations like commits and purges, older Polaris versions built their own FileIO client from those settings unless the catalog's storage config overrode the endpoint. The result: Polaris sent storage requests, signed with operation-scoped credentials, to a host the table writer picked. Deployments that do not fully trust table writers should upgrade. The reporter was credited as vignesh a.
Ajantha Bhat also shipped the Polaris Iceberg Catalog Migrator 1.1.0 after an RC3 vote with binding support from Dmitri, Russell Spitzer, JB and Alex. Russell's check covered all 12 primary Maven artifacts and 89 tests.
Enum or free text: the R2 credential vending fight
Austen Tomek of Chicago Trading has a PR adding Cloudflare R2 support with scoped credential vending. The discussion reached a clean disagreement. Austen explained why endpoint matching cannot identify the storage vendor: MinIO and other self-hosted stores use arbitrary hostnames, and a failed match silently falls back to STS. So the PR adds an explicit credentialVendingMechanism field to the Management API.
Dmitri Bourlatchkov wants that field to be free text with runtime validation, so downstream builds of Polaris can add custom mechanisms. Yufei Gu wants an enum, arguing that a spec is a contract and clients should not depend on a specific deployment's extensions. When Dmitri asked for lazy consensus, Yufei said there was none. JB stepped in on September 30 and asked everyone to slow down and align. He pointed out that the field lives in the Management API for catalog administrators. Iceberg REST clients such as Spark and Trino never see it. They only consume the standard vended credentials. That observation lowers the stakes for engine interop, but the question of how extensible the Polaris Management API should be remains open.
Persistence: DynamoDB and read replicas
Yuewei Zhou proposed an optional DynamoDB persistence backend built as a new adapter under AtomicOperationMetaStoreManager, the same shape the relational JDBC backend uses. Robert Stupp prefers a different home. He wants DynamoDB as another backend inside the existing NoSQL persistence framework, where a compare-and-set on a reference is the moment a validated catalog change becomes visible as one consistent state. He also has a DynamoDB implementation in his own fork. Dennis Huo asked the community to use DynamoDB as a chance for an apples-to-apples comparison between the two persistence approaches. JB sided with Robert. He noted that the atomic manager implicitly assumes relational multi-entity transactions, and forcing DynamoDB into that shape means orchestrating TransactWriteItems for multi-entity writes. Yuewei offered to help on whichever path moves forward.
A second persistence thread ended with a request for data. Yong Zheng wants Polaris to route read-only requests to a database read replica when a custom header is present. Prithvi S warned that routing by HTTP verb risks double ingestion if a read hits the replica before a commit is visible and the client retries. Robert Stupp read the Medium post behind the proposal and found 15 Polaris pods serving about 400 requests per second, which is under 27 per pod. That does not look like connection saturation to him. Yufei asked for baseline numbers before any design work: request rates, tables per namespace, latency targets. Yong agreed to run benchmarks with the polaris-tools harness over the coming weeks. JB added that pool multiplexers such as PgBouncer or RDS Proxy belong in that comparison.
Metrics, tags and federation
Dmitri Bourlatchkov flagged Management API changes inside PR #4115, the long-running work on REST endpoints for table metrics and events. The PR adds a TABLE_READ_METRICS grant at table, namespace and catalog level. Reading reports needs that grant, TABLE_FULL_METADATA or CATALOG_MANAGE_CONTENT. Sending scan and commit reports stays tied to data read and write privileges. Yufei reminded the list that the PR also introduces a new REST spec for metrics consumption. Dmitri proposed merging on October 1, and JB agreed, with the JDBC query SPI in PR #4756 as the next step.
EJ Wang reported progress on the Tag spec. All four planned PRs are open: the API contract in #5366, definition CRUD, assignment writes and storage, and reads with inheritance and reverse lookup. Dmitri supports merging the API PR once open review threads close. EJ also built a proof of concept for archiving design proposals as Markdown. Authors keep collaborating in Google Docs, add a version label and a link to one YAML manifest, and the tool captures the doc as Markdown with images for the repository.
David Chaava wants fail-fast validation for vendor-specific Iceberg REST federation settings, starting with BigLake. Dmitri, JB and Prithvi all asked that PolarisAdminService stay vendor-neutral. David's revised plan uses a CDI-discovered validator SPI with a BigLake implementation. Srinivas Rishindra posted the first PR for multiple storage configurations per catalog, and Travis Bowen weighed in on where the new Open Sharing APIs should live, leaning toward the existing /api/management path.
Code health
Two cleanup threads resolved. Yufei Gu found the authentication flow confusing while reviewing another PR. Two augmentors attached a credential that a third consumed, and the ordering only held because of a code comment the compiler cannot check. Alex Dutra proposed merging the OIDC pieces into a single OidcIdentityPreparer that the authenticating augmentor calls directly. Prithvi agreed, Alex opened PR #5612 on September 25, and it merged on September 28.
The long argument over schema upgrade instructions also ended. Yufei argued that SQL snippets in the upgrade guide break for anyone who starts from an older schema version, and that the versioned schema files are the only source of truth. Alex disagreed but removed every migration snippet from PR #5349 to unblock it. The page now explains how to diff the schema files. Prithvi asked for some SQL guidance to return later, since a schema diff shows the end state but not how to migrate tables that already hold data. The PR has approvals from Dmitri, Yufei and Yong Zheng.
Apache Arrow
Dropping Intel Macs
The macOS Intel thread brought numbers. Joris Van den Bossche pulled 30 days of PyPI downloads: x86_64 made up 6.6 percent of PyArrow's macOS downloads, 57,121 against 807,414 for arm64. Sutou Kouhei reproduced the check with a public ClickHouse query against the PyPI dataset and voted +1. Wes McKinney offered his 2019 Intel Mac Pro for testing commits to main. Kurtis Wright asked whether Arrow has any policy guaranteeing a number of releases before a platform goes away. Raúl Cumplido said no, and that such a promise is not realistic when Arrow depends on package managers and CI platforms with their own deprecation schedules. Homebrew already dropped Intel support without naming a last version.
Modularity and schemas
Raúl Cumplido also opened a discussion on PyArrow modularity. Arrow C++ is already split into a dozen libraries, from Core and Compute through Acero, Dataset, Flight, Flight SQL ODBC and S3. GCS and Azure still sit in core libarrow, and he wants them moved out the way S3 and the AWS SDK were. The goal is PyArrow installs that pull only the pieces a user needs, which matters for container images and serverless functions where package size has a cost.
Dewey Dunnington endorsed Kent's revised proposal for a JSON representation of Arrow schemas. David Lee asked for union type support and shared the alias-based JSON format he already uses to define data contracts between processes. A standard JSON schema form gives config files, APIs and agents a way to describe Arrow data without a binary IPC message.
Arrow Rust 60.0.0
Andrew Lamb announced Arrow Rust 60.0.0 on September 29. The release adds new Parquet encoding support, continues a sustained push to remove panics from the codebase, and lands kernel performance work plus fixes across decimals, take and filter, and FFI. The panic removal effort deserves a mention. A library that returns errors instead of aborting the process is easier to embed in long-running query engines and services, which is where arrow-rs lives inside DataFusion, Comet and a long list of downstream projects.
Apache Parquet
The vector type argument
The busiest Parquet thread of the week, at 30 messages, was about vectors. Rok Mihevc opened PR #624 in parquet-format for a LIST-based VECTOR logical type. It separates the logical contract from physical layout changes. Every non-null vector has exactly num_elements. Elements are non-null and finite. Elements are numeric or boolean primitives, with no nesting.
Antoine Pitrou objected to the restrictions. Parquet is a general-purpose format, he said, and NumPy produces vectors with NaN and infinity all the time. Readers can refuse what they do not support, but the spec should not. Alkis Evlogimenos answered with a survey. Postgres, Oracle, MySQL, SQL Server, Milvus, Qdrant and Pinecone all disallow nulls, NaN and infinity in vectors. DuckDB and ClickHouse allow them, but only because they store vectors as plain float arrays and switch off vector features when bad values appear. Daniel Weeks argued that a vector has mathematical meaning, with dimension, magnitude and direction, and non-finite values break it. Russell Spitzer asked for a concrete tool that produces non-finite vectors and a downstream engine that uses them productively.
Will Edwards reframed it. He had first assumed the goal was general ML tensors. After the sync calls he understood the immediate need was fixed-length embeddings for nearest-neighbor search, and in that context the constraints make sense. He also noted that embedding APIs do sometimes return NaN on overflow, which is exactly the kind of bad data a strict type catches. Andrew Lamb suggested a narrower name such as FINITE_VECTOR or VECTOR_EMBEDDING if the name is the real problem. Gang Wu proposed treating a general FixedSizeList as the physical layout and the specialized vector as a separate logical contract. Philipp Fischbeck and Divjot Arora both want to proceed with the proposal as written. Divjot added a practical point: statistics cannot reliably detect non-finite values, so a type-level guarantee is the only way a reader can trust the data without scanning it.
The practical outcome is taking shape. A restricted vector type for embeddings, with a possibly more specific name, and a separate track for general multidimensional arrays that Rok still wants to pursue. For lakehouse teams storing embeddings next to their tables, this is the thread that decides whether engines can index Parquet vector columns without validating every value first.
Encodings: PFOR, FastLanes, ALP and FSST
Prateek Gaur posted hard numbers in the PFOR encoding thread. Antoine had asked whether PFOR's delta mode can reuse Parquet's DELTA_BINARY_PACKED kernel. Alkis asked whether a faster DBP decoder closes the gap. Prateek measured three setups on AWS Graviton 4. On 18 datasets where delta mode helps, PFOR's own patched-delta layout reached a 10.07 compression ratio and 9.85 GB/s decode. A DBP stream inside each PFOR vector reached 6.54 and 3.76 GB/s. A tuned DBP decoder with SIMD prefix sums climbed to 5.96 GB/s but kept the 6.54 ratio. His summary: the patched-delta layout is 35 percent smaller and 1.65 times faster to decode on the datasets where delta matters. The gap is structural. PFOR uses one bit width per 1,024-value vector, while DBP splits the same vector into miniblocks with their own widths.
In the parallel FastLanes thread, Prateek and Alkis agreed to decouple FastLanes from PFOR. The initial PFOR proposal keeps sequential bit packing, and a layout discriminator leaves room to add FastLanes or interleaved layouts later. The thread also produced the sharpest exchange of the week. Antoine told Prateek that his design document and PR descriptions read like AI-generated text and asked him to rewrite them as a person talking to people. Prateek trimmed the document and cut two sections. The Parquet community is receptive to benchmark-driven proposals. It also expects authors to explain their own work.
On the rest of the encoding roadmap, the ALP blog post went through community review and Andrew Lamb described ALP as out the door. He then turned to the FSST proposal and suggested an optional inline symbol table so FSST also applies to dictionary pages. ALP targets floating point columns. FSST targets strings. Together with PFOR for integers, Parquet is on track to cover the main column types with modern lightweight encodings.
Types and the edges of the format
Micah Kornfield pushed back on two type proposals with the same argument: Parquet is a file format, not a table format. Thomas Kissinger wants an extensible decimal floating-point type with a write-time canonicalization contract. Micah answered that an empty table means no Parquet file exists, so there is no Parquet schema to coordinate independent writers. That coordination belongs to a layer above. Micah also raised concerns about a proposed WIRE logical type for serialized messages such as protobuf, citing the edge cases in lining up type systems for shredding, from recursive schemas to well-known types.
Steve Loughran argued in the format version thread that readers must fail on an unsupported version. Pretending to support a version you do not understand is a pretense, he said, and that rule is a strong reason not to bump the version without a compelling payoff. Adrian Garcia Badaracco proposed classifying distinct_count as inexact, since the spec never says whether writers are allowed to estimate it or whether readers can trust it as exact. And Danica Fine invited the community to Lakehouse Day EU on October 10 in Glasgow, co-located with Community Over Code.
Apache DataFusion
Comet 1.1.0 and the token budget
Andy Grove opened the Comet 1.1.0 RC1 vote on September 28. L. C. Hsieh, Marko Milenković and Andrew Lamb each gave a binding +1 within a day. Then Andy posted that his deeper verification had already found several regressions. He is tracking them in issue #6402 and said the full check will take a few days because he does not have an unlimited token budget. That line tells you how Andy verifies releases now: with agents that run through test matrices, where the constraint is API spend rather than hours in the day. Watch for a new candidate.
Ballista wants to ship without waiting
Andy also proposed releasing Ballista's Python client separately from the Rust crates. Ballista moved to DataFusion 55 five weeks ago and still cannot release, because the bundled Python client needs datafusion-python 55.0.0, whose vote has not started. Over the last eleven DataFusion releases, datafusion-python followed DataFusion by a median of 31 days and once by more than seven weeks. More than 300 commits sit on Ballista's main branch, including fixes for wrong results from distributed NOT IN, executors being OOM-killed on heavy queries, and jobs hanging when all executors disappear. Andy wants feedback on issue #2511, especially from Python client users.
Memory accounting
Andrew Lamb's sync notes from September 23 centered on memory. Raz from Flarion reported a string of memory issues where the memory pool's accounting did not match what operators actually used. One example was a Filter allocation that the pool never saw. Comet has hit the same gap between pool accounting and process memory. An earlier allocator-based memory test was reverted, and Andy agreed to find or create a ticket to coordinate the effort. For anyone running DataFusion inside a service with a hard memory limit, accurate pool accounting is the difference between graceful spilling and an OOM kill.
Apache Ossie (incubating)
A standard REST API for semantic models
Jean-Baptiste Onofré proposed a standard REST API for producing, consuming and orchestrating Ossie models, starting as a rest/openapi.yaml in the main repository. The response showed real vendor pull. Sebastien Gandon said Qlik has started building an Ossie-native API to import and export its data product semantic models, and he asked for a dedicated Content-Type header. Yufei Gu asked whether it builds on the Polaris Ossie REST design, which JB co-authored. JB said yes, and that the goal is for Polaris, Qlik and others to implement one standard API so clients work without knowing the server.
Viktor Kessler asked the group to agree on the problem first. He saw two scopes: a registry for exchanging and discovering models, and governed semantics, where a service guarantees that a model is bound to the physical tables it references. JB agreed and proposed two API specs, one for model management and one for catalog association. Vinoth Chandar asked that files and storage locations be first-class alongside table formats. Jakub Moravec asked how an importer knows a model changed. JB proposed keeping the core spec plain REST, with explicit versioning, content hashes and HTTP conditional requests using ETags, instead of push protocols.
This is the most interesting cross-project thread of the week. Iceberg standardized how engines find tables. Polaris implemented it. Ossie is now trying to standardize how tools find and exchange the business meaning that sits on top of those tables, and Polaris 1.8.0 already added the privileges to govern semantic models.
Converters move out, and a first release takes shape
The converter scope discussion reached a decision. Russell Spitzer favored separating the spec from implementations, the way Parquet splits parquet-format from its libraries. Markus Weimer cared less about repository count and more about keeping everything inside the ASF, since that makes adoption easier for downstream users. Kurt Stirewalt laid out what stays in the main repo: the spec, the typed object model with parsing and serialization, the validator, and the parser for the ontology formula language. JB then created the ossie-converters repository on September 29 and asked contributors to pause converter work while he moves the code.
JB also laid out the path to the first incubating releases: freeze converters, move them, and bump the spec version to 0.3.0-incubating after the recent flat-document change. Kyoung Min Kim did the arithmetic. The 0.2.0.dev0 string appears in 111 files, 96 of them under converters, so the bump is one change if it lands before the move and a coordinated two-repo change after. Kyoung Min also confirmed that the license header check now runs on every PR and the tree is clean.
What does orders.amount mean?
damian waldron asked a question the spec does not settle. In a portable expression like SUM(orders.amount), is amount a field declared in the dataset, or a physical column in the warehouse? He surveyed ten converters and found three different answers. JB said expressions and relationships must reference logical fields, because a semantic layer that reaches past its own fields to physical columns breaks encapsulation. Julian Hyde objected to any reading that locks layers into rigid tiers, since that makes Ossie less composable. Justin Talbot framed it as the SQL view case: a view's query references its immediate inputs, and dataset fields should work the same way. JB clarified that he meant lexical scoping, not a restriction on composition, and the group converged.
Two spec proposals depend on that answer. Chris Eubank narrowed the dataset-scoped metrics proposal in PR #343 so it can reach a vote without settling every open semantic question. JB strongly backed Justin's Relational Query Interface in PR #354, which grounds Layer 2 in the Measures in SQL model so BI tools that generate their own SQL still get correct metrics across fan-out and chasm traps. Justin also posted a short overview of the four spec layers in PR #452 for newcomers.
Votes, validation and structure
Chris Eubank called a vote to add the SQL-standard FILTER (WHERE ...) aggregate modifier to the expression language, after incorporating JB's feedback. The change is additive and avoids the edge cases of CASE-based workarounds for AVG, MIN, MAX and STDDEV. JB voted +1.
Kyoung Min Kim noticed that validate.py stops at the JSON schema for ontology documents because every semantic check assumes a datasets key. He will add ontology checks for built-in concepts, cycle detection on extends, and a CI step, with JB reviewing. Shao Xie's PR #458 splits ontology, semantic model and mapping into separate documents connected by references, so a slow-moving ontology can map to many semantic models. Yufei asked for more voices on deprecating embedded models. On GitHub, MikeCarlo proposed storing custom extension data as structured objects instead of opaque JSON strings, which helps the Power BI TMDL converter keep more metadata. Contributors also opened discussions on metric additivity and on recording who certified a metric definition.
An outsider builds Ossie on ClickHouse
The most useful message on the Ossie list came from a newcomer. Aleksandr Kosachev introduced himself as a backend architect who, six months ago, barely knew what a semantic layer was. He kept seeing teams wire AI assistants to databases and get confident, wrong answers. So he built ossie-clickhouse from the spec alone. It loads an Ossie document with the apache-ossie package, translates expressions to ClickHouse SQL with SQLGlot, plans one SELECT per question, and serves the model to agents over MCP. The TPC-DS example model runs end to end.
His list of places where he had to guess is a free conformance review. Relationship cardinality is not in the spec, so he only joins when the target columns form a primary or unique key. ClickHouse tables that keep change history need FINAL on read, which he exposed through a custom extension. Function semantics are open: whether REGEXP_REPLACE replaces the first match or all, whether DAYOFWEEK is ISO, and what 1/0 returns. He pinned his choices in tests until the compliance suite lands. Two of those gaps already have active threads this week. That is how a spec should improve.
Cross-Project Themes
Tighter contracts, fewer surprises. Iceberg removed a write path from v4. Parquet is leaning toward a vector type that rejects values most vector stores reject anyway. Ossie is defining whether a name means a field or a column. Polaris is arguing about whether a config field should be an enum. The shared direction is specs that make fewer promises and make them precisely, so a reader can trust data without re-validating it. The pushback is also consistent. Antoine Pitrou in Parquet, Daniel Weeks on the delegation header and Dmitri Bourlatchkov on R2 each defended flexibility for implementers. The best outcomes this week kept both. Polaris keeps its fail-fast behavior as an implementation choice while the Iceberg spec stays loose. Parquet looks headed toward a strict embedding type plus a separate general array track.
Agents are now participants in the process. Kurtis Wright found the Iceberg license gap with Claude. Kevin Liu pointed an agent at the repository to hunt correctness bugs. Andy Grove's Comet verification is gated by token budget. Gaspard Merten wrote his PyIceberg change with Claude. The communities are also setting norms. Contributors label AI-assisted votes, and Antoine asked a Parquet author to rewrite AI-sounding prose in his own words. Agents also showed up as consumers. Andrei Tserakhau argued for SQL-visible Iceberg labels because agents explore warehouses through SQL, and Aleksandr Kosachev built an Ossie server specifically to stop agents from giving wrong answers.
The semantic layer is joining the catalog layer. Polaris 1.8.0 added privileges for semantic models. Ossie proposed a REST API modeled on the Polaris design and split out a catalog association spec. Iceberg debated catalog-owned labels. The lakehouse stack now has open standards for files, tables and catalogs, and the next layer up, what the numbers mean, is getting the same treatment on the same lists with many of the same people.
Release engineering is a security surface. Iceberg RC1 failed on bundled license gaps. Polaris rc1 failed on a license header. Confluent is scanning the Iceberg Kafka connector. Polaris shipped a CVE fix for server-side FileIO endpoints. The Iceberg test backend moved from MinIO to RustFS after CI breakage. None of this is glamorous. All of it is the work that lets enterprises run these projects in production.
Looking Ahead
Watch for the Iceberg 1.12.0 announcement and the equality delete spec merge, then the October 8 sync on File and Vector types. Polaris plans to merge the table metrics PR on October 1, and the R2 enum question needs a resolution before that PR lands. Parquet's vector type should get a naming decision and a path to a vote. DataFusion Comet needs a new release candidate once Andy Grove's regression list is triaged. Ossie will move converters into their new repository and prepare its first incubating release at spec version 0.3.0. And if you are in Europe, Lakehouse Day EU in Glasgow on October 10 is the next place to see many of these contributors in one room.
Keep Learning
If this digest helps you follow the open lakehouse, my books go deeper on every layer covered here, from Apache Iceberg table design and Apache Polaris catalogs to Arrow, Parquet and building agentic analytics on open data. Find the full catalog at books.alexmerced.com.
Top comments (0)