<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Zephico Technologies</title>
    <description>The latest articles on DEV Community by Zephico Technologies (@zephico).</description>
    <link>https://dev.to/zephico</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4029027%2F534e6458-b5b1-498a-8e07-42f5db7432e2.png</url>
      <title>DEV Community: Zephico Technologies</title>
      <link>https://dev.to/zephico</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/zephico"/>
    <language>en</language>
    <item>
      <title>Domo: what it actually does, and where it earns its keep</title>
      <dc:creator>Zephico Technologies</dc:creator>
      <pubDate>Sat, 05 Sep 2026 15:08:41 +0000</pubDate>
      <link>https://dev.to/zephico/domo-what-it-actually-does-and-where-it-earns-its-keep-2nbh</link>
      <guid>https://dev.to/zephico/domo-what-it-actually-does-and-where-it-earns-its-keep-2nbh</guid>
      <description>&lt;p&gt;Most BI conversations start with a chart tool and end with a data-engineering project — someone needs a dashboard, and getting there means pipelines, a warehouse, a semantic layer, then finally the visualization. Domo's pitch is that it collapses that whole chain into one product: connectors, transformation, storage, dashboards and now AI-assisted analysis, all under one login. That's genuinely useful for the right team, and genuinely the wrong shape for another. Here's the honest breakdown.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Domo actually bundles
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Connectors, not integration projects.&lt;/strong&gt; Domo ships with well over a thousand prebuilt connectors — Salesforce, NetSuite, Google Sheets, ad platforms, databases, other SaaS tools. For a business team that would otherwise wait on an engineer to build an extract job, pointing Domo at a source and getting a table in minutes is the actual selling point, not a footnote.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Magic ETL as a visual pipeline builder.&lt;/strong&gt; Instead of writing SQL or Python, you drag transformation steps — join, filter, aggregate, pivot — onto a canvas. It's genuinely low-code, which means an analyst who isn't an engineer can build and maintain a real pipeline, not just a report.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Dashboards built for executives, not just analysts.&lt;/strong&gt; Domo's cards and dashboards are polished by default, work on mobile without extra effort, and support drill-down and alerting out of the box. "Domo Everywhere" lets you embed or white-label dashboards into a customer-facing product, which is a real differentiator if you're selling analytics as part of your own SaaS rather than only consuming it internally.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Domo.AI on top.&lt;/strong&gt; Natural-language querying against your data, auto-generated summaries, and agent-style workflows are now part of the platform rather than a bolt-on — useful for self-serve questions that would otherwise turn into a ticket to the analytics team.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where that's the right call
&lt;/h2&gt;

&lt;p&gt;Domo is at its best for mid-market and departmental analytics: marketing, sales ops, finance teams that have real data scattered across a dozen SaaS tools and no dedicated data engineering function to unify it. If the alternative is "nobody looks at this data because getting it into one place takes a sprint," Domo's speed-to-first-dashboard is the whole business case, and it's a legitimate one. It's also a strong fit when the deliverable isn't an internal dashboard at all but an analytics feature you're shipping to your own customers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it runs out of road
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Cost scales with consumption, and that surprises people.&lt;/strong&gt; Domo's pricing has historically tracked rows and data volume rather than a flat per-seat fee. That's fine at departmental scale and can get expensive fast once a company tries to route serious data volume — event-level logs, high-frequency transactional data — through it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The pipelines aren't portable.&lt;/strong&gt; Magic ETL logic lives inside Domo's proprietary model. There's no equivalent of exporting a dbt project or a set of Delta tables and running it somewhere else — if you outgrow Domo, you're rebuilding the transformation layer, not migrating it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It's a presentation and light-transform layer, not a governed data platform.&lt;/strong&gt; Domo can query and join what you point it at, but it isn't a substitute for a proper lakehouse with lineage, access control at the table level, and a single source of truth other systems (not just dashboards) can read from. Teams that start on Domo because it's fast often hit a second project a year or two later: standing up the governed platform Domo should have been sitting on top of, rather than acting as, all along.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual decision
&lt;/h2&gt;

&lt;p&gt;If the honest answer to "where does our data live today" is "scattered across SaaS tools, with no engineering team dedicated to unifying it," Domo will get a real dashboard in front of the business faster than any alternative, and that speed has value. If the honest answer is "we already have — or need — a governed warehouse or lakehouse that other systems depend on," treat Domo as a front-end option on top of that platform, not a replacement for building it.&lt;/p&gt;

&lt;p&gt;We build the layer underneath this decision for clients — &lt;a href="https://zephico.com/services/data-analytics-ai" rel="noopener noreferrer"&gt;semantic layers and executive dashboards&lt;/a&gt; on a governed data platform, whether the presentation layer on top ends up being Domo, Power BI, or something custom. If you're trying to work out which side of that line your team is on, &lt;a href="https://zephico.com/contact" rel="noopener noreferrer"&gt;talk to us&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://zephico.com/blog/domo-business-intelligence-platform" rel="noopener noreferrer"&gt;Zephico blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>domo</category>
      <category>businessintelligence</category>
      <category>analytics</category>
    </item>
    <item>
      <title>LLM features: the distance between the demo and production</title>
      <dc:creator>Zephico Technologies</dc:creator>
      <pubDate>Tue, 01 Sep 2026 09:27:24 +0000</pubDate>
      <link>https://dev.to/zephico/llm-features-the-distance-between-the-demo-and-production-3id5</link>
      <guid>https://dev.to/zephico/llm-features-the-distance-between-the-demo-and-production-3id5</guid>
      <description>&lt;p&gt;Every company has now seen the demo: someone wires a model to internal documents, asks it a question, and the room goes quiet. The demo takes a week. The gap between that and a feature you'd put in front of customers is the actual project, and &lt;a href="https://zephico.com/services/data-analytics-ai" rel="noopener noreferrer"&gt;we build in that gap&lt;/a&gt; — so here's an honest map of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the demo lies
&lt;/h2&gt;

&lt;p&gt;A demo answers the questions its builder asks it, on the documents they chose, judged by impression. Production answers whatever users actually type, over your full messy corpus, judged by consequences. Every hard problem lives in that difference.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Retrieval is most of the product.&lt;/strong&gt; For anything grounded in your data, answer quality is mostly determined before the model sees the question — by how documents are chunked, indexed and searched. Garbage retrieval with a frontier model still produces confident garbage. This is data engineering, not prompt magic, and it's where most of the build time goes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Evals are the tests you can't skip.&lt;/strong&gt; With normal code you write tests; with a probabilistic system you need evaluations — a graded set of real questions with known-good answers, run on every change. Teams without evals ship on vibes, and every prompt tweak is a gamble on regressions they can't see. Building the first eval set is unglamorous work with the highest ROI in the project.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Failure needs a designed path.&lt;/strong&gt; The model will sometimes be wrong. The product question is what happens then: Can it say "I don't know"? Does it cite sources so users can verify? Is there a human escalation for consequential actions? Features fail in production not because the model errs, but because nobody designed for the error.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cost and latency are product constraints.&lt;/strong&gt; Tokens are cheap; tokens times retrieval context times every user times every day is a real bill, and multi-step chains can turn a "fast" model into an eight-second wait. Decisions like caching, model tiering (small model for easy cases, big model for hard ones) and context budgets belong in the design, not the postmortem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Security gets a new attack surface.&lt;/strong&gt; User input is now instructions. Prompt injection — including through the documents you retrieve — plus PII flowing into prompts and logs means access control has to live &lt;em&gt;outside&lt;/em&gt; the model. The model must never be the thing enforcing who can see what.&lt;/p&gt;

&lt;h2&gt;
  
  
  The staging that works
&lt;/h2&gt;

&lt;p&gt;Ship it in this order and each step pays for the next: &lt;strong&gt;internal assistant first&lt;/strong&gt; (support team, not customers — cheap feedback, contained failures), &lt;strong&gt;evals before broadening&lt;/strong&gt; (built from real internal usage), &lt;strong&gt;customer-facing with sources and escape hatches&lt;/strong&gt;, and only then autonomy over consequential actions, if ever. Teams that invert this — customer-facing agent first — generate incident reports, not product.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to skip the LLM entirely
&lt;/h2&gt;

&lt;p&gt;If the workflow is deterministic — lookups, calculations, routing on known fields — conventional software is cheaper, faster and doesn't hallucinate. Some of the best consulting sentences we deliver start with "you don't need a model for this." An unglamorous &lt;a href="https://zephico.com/services/data-analytics-ai" rel="noopener noreferrer"&gt;dashboard on trusted data&lt;/a&gt; beats an oracle nobody trusts.&lt;/p&gt;

&lt;p&gt;If you have a use case and a corpus, &lt;a href="https://zephico.com/contact" rel="noopener noreferrer"&gt;we'll tell you which side of that line you're on&lt;/a&gt; — and if it's the model side, build the retrieval and eval foundations that make it real.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://zephico.com/blog/llm-features-demo-to-production" rel="noopener noreferrer"&gt;Zephico blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>dataengineering</category>
    </item>
    <item>
      <title>Databricks or Snowflake? How mid-size companies should choose</title>
      <dc:creator>Zephico Technologies</dc:creator>
      <pubDate>Tue, 01 Sep 2026 09:27:02 +0000</pubDate>
      <link>https://dev.to/zephico/databricks-or-snowflake-how-mid-size-companies-should-choose-46jj</link>
      <guid>https://dev.to/zephico/databricks-or-snowflake-how-mid-size-companies-should-choose-46jj</guid>
      <description>&lt;p&gt;Full disclosure up front: we're a &lt;a href="https://zephico.com/services/data-engineering-databricks" rel="noopener noreferrer"&gt;Databricks shop&lt;/a&gt; — our engineers are certified on it and it's where we build. That's exactly why this comparison is worth writing honestly: we've watched companies choose each platform for good and bad reasons, and the bad reasons are always the same.&lt;/p&gt;

&lt;h2&gt;
  
  
  The convergence is real, the centers of gravity aren't
&lt;/h2&gt;

&lt;p&gt;On paper the platforms now do each other's jobs: Snowflake runs Python and ML workloads, Databricks ships a SQL warehouse with BI-grade performance. But each platform still has a center of gravity, and you will live near it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Snowflake's center is SQL analytics.&lt;/strong&gt; If your world is structured data, dashboards and analysts who live in SQL, Snowflake is genuinely excellent: near-zero administration, predictable behavior, warehouses your analysts cannot break.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Databricks' center is engineering on data.&lt;/strong&gt; Streaming pipelines, unstructured data, machine learning, and transformations complex enough to deserve version control and tests. If your roadmap includes "then we build ML/AI features on this," you'll end up doing engineering work — and Databricks is where that work is native rather than bolted on.&lt;/p&gt;

&lt;h2&gt;
  
  
  The questions that actually decide it
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Who works on the data?&lt;/strong&gt; Count heads. Mostly SQL analysts → Snowflake will make them productive on day one. Data engineers writing Python → Databricks gives them a real development environment instead of a SQL console with extensions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What's the ugliest data you'll touch in two years?&lt;/strong&gt; If the answer is "CSV exports and a CRM," either platform is fine. If it's event streams, documents, images or model training data, be honest about that now — migrating platforms later costs more than choosing right.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How does the bill grow?&lt;/strong&gt; Both are consumption-billed, and both will punish you for ungoverned usage — Snowflake through warehouses someone left running on autopilot, Databricks through clusters someone sized by vibes. The difference we see in practice: Databricks gives engineers more knobs, which means more savings &lt;em&gt;if&lt;/em&gt; someone owns cost engineering, and more waste if nobody does. Budget for governance either way — on Databricks that means &lt;a href="https://zephico.com/blog/unity-catalog-migration-what-it-takes" rel="noopener noreferrer"&gt;Unity Catalog&lt;/a&gt;, which is no longer optional.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How much lock-in can you tolerate?&lt;/strong&gt; Databricks stores data in open formats (Delta Lake, now interoperable with Iceberg) on your own cloud storage — your data remains yours to point other engines at. Snowflake has moved toward open formats too, but its gravity is still data living inside Snowflake. For a company that expects to renegotiate in five years, open storage is a real bargaining chip.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bad reasons we keep seeing
&lt;/h2&gt;

&lt;p&gt;Choosing Databricks because "we might do AI someday" with no engineer to build it — you'll pay a complexity tax for an option you never exercise. Choosing Snowflake because the finance team liked the demo dashboard — and then hiring data engineers who spend their lives working around it. Choose for the team you'll actually have.&lt;/p&gt;

&lt;h2&gt;
  
  
  A structure that works either way
&lt;/h2&gt;

&lt;p&gt;Whichever engine you pick, the architecture that keeps you sane is the same: raw data preserved untouched, cleaned and conformed layers on top, business-ready tables at the end. We wrote a &lt;a href="https://zephico.com/blog/medallion-architecture-databricks-practical-guide" rel="noopener noreferrer"&gt;practical guide to the medallion architecture&lt;/a&gt; that covers how we structure it.&lt;/p&gt;

&lt;p&gt;If you're weighing this decision with real workloads on the table, &lt;a href="https://zephico.com/contact" rel="noopener noreferrer"&gt;we're happy to look at them with you&lt;/a&gt; — including the cases where the honest answer is Snowflake.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://zephico.com/blog/databricks-vs-snowflake-mid-size-companies" rel="noopener noreferrer"&gt;Zephico blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>databricks</category>
      <category>dataengineering</category>
    </item>
    <item>
      <title>Shipping a data app on the lakehouse: Databricks Apps in practice</title>
      <dc:creator>Zephico Technologies</dc:creator>
      <pubDate>Thu, 27 Aug 2026 22:47:07 +0000</pubDate>
      <link>https://dev.to/zephico/shipping-a-data-app-on-the-lakehouse-databricks-apps-in-practice-57i</link>
      <guid>https://dev.to/zephico/shipping-a-data-app-on-the-lakehouse-databricks-apps-in-practice-57i</guid>
      <description>&lt;p&gt;Most internal tools that read lakehouse data are built the same way: stand up a small app somewhere (Retool, a hand-rolled Streamlit deployment, a Flask service on its own host), then wire it up to Databricks with a service-account token and hope nobody rotates credentials at the wrong time. Databricks Apps removes the "somewhere" — it hosts the app inside the workspace itself, running Streamlit, Dash, Gradio, or a custom Flask app, with workspace identity and Unity Catalog access available natively instead of bolted on. Here's what that's actually good for, and where it isn't the right tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's actually different, not just convenient
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Auth is the workspace's, not a separate login you build.&lt;/strong&gt; An app deployed on Databricks Apps authenticates users through the workspace's own identity — OAuth against Databricks accounts — rather than a bespoke login screen and a separate user table to maintain. For an internal tool, that's not a minor convenience: it's one fewer identity system to keep in sync with who actually has access to what.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Data access flows through Unity Catalog governance, not a static token.&lt;/strong&gt; This is the detail worth getting right early, because it determines whether the app's data access is actually governed or only looks like it is: an app can query as its own service-principal identity, or run queries on behalf of the logged-in viewer, and the choice changes everything. Run as the app's own identity and every viewer sees whatever that identity can see — UC's row filters and column masks apply once, at the app level, not per person. Run on behalf of the viewer and UC's row filters, column masks, and grants apply exactly as they would if that person queried the table directly in a SQL warehouse. For anything with real access-control requirements — an app different teams use to see different slices of the same table — that second mode is the one that actually enforces the governance model you already built in Unity Catalog, rather than reimplementing an approximation of it in application code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No second infrastructure stack for a purely internal, lakehouse-native tool.&lt;/strong&gt; A data-quality review UI, an approval workflow that writes back to a Delta table, a lightweight ops dashboard with a form — tools like this traditionally need their own hosting, their own database connection management, and their own deployment pipeline, none of which has anything to do with the tool's actual job. Databricks Apps collapses that into "the app lives where the data already lives," which is a genuine reduction in what has to be operated, not just a nicer developer experience.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it fits and where it doesn't
&lt;/h2&gt;

&lt;p&gt;The sweet spot is internal, lakehouse-native tools: data-quality review and approval UIs, ops dashboards with write-back to governed tables, lightweight self-service tools for teams that already work inside Databricks. If the tool's entire job is reading and writing Unity Catalog data with per-user permissions that should already exist, building it as a Databricks App means the governance model you built for SQL access carries straight through to the app, instead of being re-specified in a second system.&lt;/p&gt;

&lt;p&gt;It's a weaker fit once a tool needs to integrate many systems that aren't Databricks — a support tool pulling from a CRM, a ticketing system, and three SaaS APIs alongside the lakehouse — where an integration-first platform like Retool's breadth of prebuilt connectors is doing real work that a lakehouse-native app framework isn't built for. It's also not the right choice for high-traffic, customer-facing applications: it's built for internal tooling with workspace-scoped auth, not a public app needing its own scaling and deployment pipeline outside the workspace boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  The practical takeaway
&lt;/h2&gt;

&lt;p&gt;Don't default to "everything internal is a Retool app" or "everything on Databricks should be a Databricks App" — ask what the tool is actually integrating. Pure lakehouse read/write with governance that should mirror Unity Catalog exactly is the case Databricks Apps was built for, and getting the run-as-viewer-versus-run-as-service-principal decision right on day one saves a governance retrofit later. Multi-system internal tools still belong on a platform built for breadth of integration — which is a large part of why we run both practices.&lt;/p&gt;

&lt;p&gt;Zephico is a &lt;a href="https://zephico.com/partners/databricks" rel="noopener noreferrer"&gt;Databricks Consulting Partner&lt;/a&gt; and also builds &lt;a href="https://zephico.com/services/internal-tools-retool" rel="noopener noreferrer"&gt;Retool and internal-tools projects&lt;/a&gt; — we help clients pick the right one for a given tool rather than defaulting to whichever we built last. If you're weighing where a new internal tool should live, &lt;a href="https://zephico.com/contact" rel="noopener noreferrer"&gt;talk to us&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://zephico.com/blog/databricks-apps-lakehouse-data-app" rel="noopener noreferrer"&gt;Zephico blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>databricks</category>
      <category>dataengineering</category>
    </item>
    <item>
      <title>Serverless Databricks: when it's cheaper and when it isn't</title>
      <dc:creator>Zephico Technologies</dc:creator>
      <pubDate>Thu, 27 Aug 2026 22:47:04 +0000</pubDate>
      <link>https://dev.to/zephico/serverless-databricks-when-its-cheaper-and-when-it-isnt-4ibl</link>
      <guid>https://dev.to/zephico/serverless-databricks-when-its-cheaper-and-when-it-isnt-4ibl</guid>
      <description>&lt;p&gt;Serverless compute on Databricks — serverless SQL warehouses, serverless compute for notebooks and jobs, serverless GPU compute for model serving — gets pitched as a strict upgrade: no cluster startup wait, no capacity planning, no idle spend from a cluster nobody remembered to shut down. All of that is true. What doesn't get said as often is that serverless carries a per-DBU price premium over the equivalent classic compute, so "it removes ops burden" and "it's cheaper" are two separate claims, and only one of them is universally true.&lt;/p&gt;

&lt;h2&gt;
  
  
  What serverless actually removes
&lt;/h2&gt;

&lt;p&gt;Databricks manages the underlying infrastructure — instance provisioning, warm pools, scaling — so there's no cluster startup delay (serverless SQL warehouses and jobs start in seconds, not the minutes a classic cluster from cold can take), no manual autoscaling configuration to get wrong, and critically, no idle classic cluster billing while nobody's running anything. For interactive and bursty workloads, that last point alone often justifies the switch: a classic all-purpose cluster left running "just in case" between an analyst's queries burns DBUs doing nothing, and serverless simply doesn't have that failure mode.&lt;/p&gt;

&lt;h2&gt;
  
  
  When it's genuinely cheaper
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Bursty, interactive workloads.&lt;/strong&gt; Ad hoc SQL and notebook work that happens in short, unpredictable bursts through the day is the clearest win — you pay for the seconds actually used, with no idle tail and no pre-provisioned capacity sitting unused between sessions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Jobs that run frequently but briefly.&lt;/strong&gt; A job that runs every ten minutes for thirty seconds pays a real cluster-startup tax on classic compute — either you keep a cluster warm (idle cost) or you eat the startup latency and its DBU cost on every run. Serverless jobs compute avoids both.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Teams that were over-provisioning to avoid the ops burden.&lt;/strong&gt; A common classic-compute pattern is sizing a cluster generously and leaving autoscaling loosely configured because nobody has time to tune it properly. Serverless removes the incentive to over-provision defensively, because there's no cluster sizing decision left to get wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  When it isn't
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Large, steady-state, long-running batch jobs.&lt;/strong&gt; A 24/7 ETL pipeline with predictable, high, sustained utilization is exactly the workload classic compute with Spot instances and committed-use discounts was built to serve cheaply. Serverless's per-DBU premium is the price of on-demand elasticity you're not using if the workload never actually varies — paying for elasticity you don't need is the single most common way serverless ends up costing more, not less.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Workloads where instance-type control matters.&lt;/strong&gt; Classic clusters let you pick instance families and rightsize for a specific workload's memory-to-compute ratio; serverless trades that control away for simplicity. If a job genuinely benefits from a specific instance shape — memory-optimized for a particular join pattern, say — classic compute keeps that lever available.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Anywhere a Reserved or Savings-Plan-style commitment already applies.&lt;/strong&gt; If you've committed to baseline capacity on classic compute (the DBU-equivalent of a cloud reserved instance), that commitment is a sunk discount working against you the moment load shifts to serverless — you're now paying for unused committed capacity &lt;em&gt;and&lt;/em&gt; the serverless premium on top, which is worse than either option alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual decision framework
&lt;/h2&gt;

&lt;p&gt;Plot your workload's utilization curve, not its label. Spiky, low-average-utilization workloads — most interactive SQL, most ad hoc notebook work, most frequent-but-short jobs — favor serverless because the premium is smaller than the idle cost it removes. Sustained-high-utilization workloads — nightly batch ETL that runs for hours at consistent load, always-on production jobs — favor classic compute with commitments, because there's no elasticity premium worth paying when the load never actually flexes. Most real Databricks accounts are a mix, and the right answer is usually "serverless for the bursty half, classic-with-commitments for the steady half," not an all-or-nothing platform choice.&lt;/p&gt;

&lt;p&gt;The one thing not to do is guess. This is exactly the kind of decision that should be made from actual &lt;code&gt;system.billing.usage&lt;/code&gt; data broken down by workload pattern, not from a vendor's default recommendation or last year's habit.&lt;/p&gt;

&lt;p&gt;Zephico is a &lt;a href="https://zephico.com/partners/databricks" rel="noopener noreferrer"&gt;Databricks Consulting Partner&lt;/a&gt;, and our &lt;a href="https://zephico.com/services/data-engineering-databricks" rel="noopener noreferrer"&gt;Databricks-certified engineers&lt;/a&gt; run cost audits that reconcile against your actual billing usage tables rather than estimating from cluster specs — including the serverless-vs-classic question specifically. If your Databricks bill has grown past the point where this decision is a rounding error, &lt;a href="https://zephico.com/contact" rel="noopener noreferrer"&gt;talk to us&lt;/a&gt; about a free DBU cost analysis.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://zephico.com/blog/serverless-databricks-cost" rel="noopener noreferrer"&gt;Zephico blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>databricks</category>
      <category>business</category>
    </item>
    <item>
      <title>Mosaic AI for a first GenAI feature: what has to be in place before model serving</title>
      <dc:creator>Zephico Technologies</dc:creator>
      <pubDate>Thu, 27 Aug 2026 22:46:28 +0000</pubDate>
      <link>https://dev.to/zephico/mosaic-ai-for-a-first-genai-feature-what-has-to-be-in-place-before-model-serving-4h50</link>
      <guid>https://dev.to/zephico/mosaic-ai-for-a-first-genai-feature-what-has-to-be-in-place-before-model-serving-4h50</guid>
      <description>&lt;p&gt;Mosaic AI is Databricks' answer to running GenAI features on top of governed lakehouse data rather than a separate ML stack bolted alongside it: Model Serving, Vector Search, Feature Serving, an Agent Framework, and MLflow's evaluation and tracing tooling, all sitting on Unity Catalog. The part teams get right quickly is standing up a serving endpoint — that's genuinely close to a few clicks or a short API call. The part that determines whether the feature is any good is everything that has to exist before the endpoint matters at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Serving is the easy part
&lt;/h2&gt;

&lt;p&gt;A Mosaic AI Model Serving endpoint, whether it's fronting a Foundation Model API (Databricks-hosted models like Llama or DBRX, billed pay-per-token), a provisioned-throughput deployment for predictable high volume, or your own fine-tuned model, is genuinely fast to stand up. That speed is exactly what makes it tempting to treat as the whole project. It isn't. An endpoint that returns fast, well-formatted, confidently wrong answers is not a shipped feature — it's a liability with good uptime.&lt;/p&gt;

&lt;h2&gt;
  
  
  What has to exist first
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Governed, retrievable data.&lt;/strong&gt; If the feature is grounded in your own data (support docs, product catalog, internal knowledge base), Mosaic AI Vector Search needs a Delta table with Change Data Feed enabled as its source, so the index stays automatically synced as the underlying table changes. This means the retrieval foundation is a Unity Catalog governance problem before it's an AI problem — the same table hygiene, access controls, and column descriptions that make any lakehouse data trustworthy are what determine whether retrieval returns the right documents.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A real evaluation set, not vibes.&lt;/strong&gt; MLflow's evaluation tooling (and Mosaic AI's Agent Evaluation on top of it) exists because a probabilistic system needs graded test cases the way deterministic code needs unit tests. Before any endpoint goes near real users, there should be a set of representative questions with known-good answers, and every prompt or retrieval change should be run against it. The Databricks review app — where domain experts label real outputs as correct or not — is the mechanism that turns "we think it's good" into a number you can track over time. Skipping this step doesn't make the project faster; it moves the discovery of quality problems from a controlled eval run to a Slack complaint from a user.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A gateway in front of the endpoint, not just the endpoint.&lt;/strong&gt; Mosaic AI Gateway sits in front of serving endpoints and handles rate limiting, usage tracking, and safety guardrails (PII detection, content filtering) centrally, rather than each application team reimplementing its own version. For a first feature this is easy to defer and easy to regret deferring — retrofitting guardrails after a feature is live and something has already gone wrong is a much worse conversation than building them in from the start.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A cost model, before volume, not after.&lt;/strong&gt; Foundation Model API pricing is per-token, and a naive retrieval-augmented prompt that stuffs in ten full documents per query has a very different cost profile than one with a tuned chunk size and a relevance threshold. Provisioned throughput trades that per-token variability for a fixed capacity cost, which is the right call once volume is predictable and wrong when it isn't. Knowing which regime you're in belongs in the design, not discovered on the first real invoice.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mistake that shows up every time
&lt;/h2&gt;

&lt;p&gt;Teams build the serving endpoint first because it's the most demoable piece, then work backward into retrieval quality and evaluation once the first round of user complaints comes in. It's cheaper, by a wide margin, to build in the order the dependencies actually require: governed data and a defined retrieval strategy first, an eval set before broad rollout, a gateway with guardrails in front of anything user-facing, and only then treat the serving endpoint as the finished piece it looks like on day one.&lt;/p&gt;

&lt;p&gt;This is the same discipline our &lt;a href="https://zephico.com/services/data-analytics-ai" rel="noopener noreferrer"&gt;data analytics and AI team&lt;/a&gt; applies regardless of which specific platform sits underneath — Mosaic AI just makes more of the plumbing native to a workspace you may already run. Zephico is a &lt;a href="https://zephico.com/partners/databricks" rel="noopener noreferrer"&gt;Databricks Consulting Partner&lt;/a&gt;; if you're scoping a first GenAI feature on Databricks and want a straight read on what has to exist before the endpoint does, &lt;a href="https://zephico.com/contact" rel="noopener noreferrer"&gt;talk to us&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://zephico.com/blog/mosaic-ai-first-genai-feature" rel="noopener noreferrer"&gt;Zephico blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>databricks</category>
      <category>ai</category>
    </item>
    <item>
      <title>What the Databricks REST API can and can't automate in a large migration</title>
      <dc:creator>Zephico Technologies</dc:creator>
      <pubDate>Thu, 27 Aug 2026 22:46:25 +0000</pubDate>
      <link>https://dev.to/zephico/what-the-databricks-rest-api-can-and-cant-automate-in-a-large-migration-513m</link>
      <guid>https://dev.to/zephico/what-the-databricks-rest-api-can-and-cant-automate-in-a-large-migration-513m</guid>
      <description>&lt;p&gt;Every large Databricks migration — Hive metastore to Unity Catalog, on-prem Spark to Databricks, or consolidating a sprawl of workspaces into a governed few — starts with the same hopeful question: "can we just script this?" The honest answer is that the Databricks REST API automates the mechanical 80% of a migration extremely well, and the remaining 20% is exactly the part a script can't safely decide for you. Knowing which is which up front is what separates a migration that finishes on schedule from one that stalls in month four.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the API automates well
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Workspace object migration.&lt;/strong&gt; The Workspace API exports and imports notebooks, folders, and Repos in bulk — this is genuinely close to a solved problem, and it's usually the first thing a migration script gets working, because it's low-risk and easy to verify (diff the exported source against the imported copy).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Jobs and cluster policies.&lt;/strong&gt; The Jobs API and Clusters API let you enumerate every job definition, schedule, and cluster policy in a source workspace and recreate them programmatically in a target one. Recreating a thousand job definitions by hand in the UI isn't a realistic option; scripting it against the API is.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Identity and access provisioning.&lt;/strong&gt; SCIM support lets you migrate users, groups, and group memberships without manually recreating every account. Combined with the Permissions API, you can replicate who has access to which jobs, clusters, and notebooks — as data, at least.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Unity Catalog object creation.&lt;/strong&gt; Catalogs, schemas, external locations, and storage credentials can all be created via API from a defined inventory, which matters because a large migration usually means creating dozens or hundreds of these objects consistently rather than one-off through the UI.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it can't automate — because the hard part isn't the API call
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Hive ACLs don't map to UC grants mechanically.&lt;/strong&gt; Hive metastore's table access control is enforced by cluster configuration — table ACL cluster settings, instance profiles scoped to S3 prefixes — not by grants stored with the table. There's no API call that reads "who could access this table under the old model" and emits the equivalent &lt;code&gt;GRANT&lt;/code&gt; statements, because the old model didn't store that information in a form a script can reliably reconstruct. Someone has to decide the new grant structure; the API only applies the decision once it's made.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;DBFS mount paths are a business-logic problem wearing a technical costume.&lt;/strong&gt; Every notebook, job, and library hardcoded to &lt;code&gt;/mnt/some-path&lt;/code&gt; needs to be found and updated to a Unity Catalog volume or external location. The API can help you &lt;em&gt;find&lt;/em&gt; every reference (grep the exported notebook source), but rewriting each one correctly requires understanding what that notebook is actually doing — a script that blindly find-and-replaces mount paths will break notebooks that reference paths conditionally or construct them dynamically.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rate limits mean "migrate everything at once" isn't a real strategy.&lt;/strong&gt; The REST API enforces per-workspace rate limits on list and get calls, and a workspace with ten thousand notebooks or a thousand job definitions will throttle a naive single-loop migration script. A migration tool that actually works at scale batches requests and backs off on 429s — this is ordinary API-client engineering, but it's exactly the kind of detail that turns a weekend script into something that silently fails halfway through a Friday-night migration window.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Correctness after translation isn't verifiable by the API.&lt;/strong&gt; Once ACLs are translated to grants and mount paths are rewritten to volumes, the API can tell you the objects were created — it cannot tell you the new permission model actually matches the intended one, or that a rewritten notebook still produces the same output. That validation is either manual spot-checking or a purpose-built comparison job, not a REST endpoint.&lt;/p&gt;

&lt;h2&gt;
  
  
  The shape of a migration that actually works
&lt;/h2&gt;

&lt;p&gt;The migrations that go smoothly split into three phases, and the API's role is different in each: an &lt;strong&gt;inventory phase&lt;/strong&gt; that's pure API reads — enumerate everything, understand scope, no changes made — a &lt;strong&gt;translation phase&lt;/strong&gt; that's mostly human judgment producing a reviewed mapping (old ACL → new grant, old mount path → new volume), and an &lt;strong&gt;apply phase&lt;/strong&gt; where the API executes that already-reviewed mapping in controlled, rate-limited batches. Skipping straight from inventory to apply, treating the translation step as something the script can infer, is where migrations that looked automatable on paper turn into months of cleanup.&lt;/p&gt;

&lt;p&gt;Zephico is a &lt;a href="https://zephico.com/partners/databricks" rel="noopener noreferrer"&gt;Databricks Consulting Partner&lt;/a&gt;, and our &lt;a href="https://zephico.com/services/data-engineering-databricks" rel="noopener noreferrer"&gt;Databricks-certified engineers&lt;/a&gt; run large workspace and Unity Catalog migrations for clients — including the judgment calls in the translation phase that no API call makes for you. If you're scoping a migration and want an honest read on what's actually scriptable versus what needs a decision first, &lt;a href="https://zephico.com/contact" rel="noopener noreferrer"&gt;talk to us&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://zephico.com/blog/databricks-rest-api-migration-automation" rel="noopener noreferrer"&gt;Zephico blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>databricks</category>
      <category>dataengineering</category>
    </item>
    <item>
      <title>Unity Catalog and open formats: Delta, Iceberg and UniForm without picking a side</title>
      <dc:creator>Zephico Technologies</dc:creator>
      <pubDate>Thu, 27 Aug 2026 22:45:45 +0000</pubDate>
      <link>https://dev.to/zephico/unity-catalog-and-open-formats-delta-iceberg-and-uniform-without-picking-a-side-9i2</link>
      <guid>https://dev.to/zephico/unity-catalog-and-open-formats-delta-iceberg-and-uniform-without-picking-a-side-9i2</guid>
      <description>&lt;p&gt;For a couple of years, "Delta or Iceberg" was a real strategic question — the kind that shaped which query engines a company could realistically use, because most engines committed to reading one table format natively and treated the other as an afterthought. That question matters a lot less than it used to. Unity Catalog now lets you write Delta and have it read as Iceberg — or vice versa — without a second copy of the data or a duplicate ETL pipeline. Here's how that actually works, and where it still isn't a total non-issue.&lt;/p&gt;

&lt;h2&gt;
  
  
  What UniForm actually does
&lt;/h2&gt;

&lt;p&gt;Delta Lake UniForm (Universal Format) generates Iceberg metadata alongside the Delta metadata your writes already produce, on top of the same underlying Parquet data files. There's no second write of the data itself — Delta and Iceberg both describe the same physical Parquet, just through two different metadata layers, kept in sync automatically on each Delta commit. An engine that only speaks Iceberg — Trino, Snowflake, Redshift Spectrum, Athena — can point at the Iceberg metadata and read the table as if it were natively Iceberg, with no awareness that Databricks wrote it as Delta.&lt;/p&gt;

&lt;p&gt;Practically, this means a table your pipelines write with Delta-native features (deletion vectors, liquid clustering, Delta's transaction log) is simultaneously queryable by any Iceberg-compatible engine your organization already runs, without a nightly export job that duplicates storage and drifts out of sync between runs. That single fact removes most of the reason companies used to maintain parallel Delta and Iceberg copies of the same data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Unity Catalog as an Iceberg REST catalog
&lt;/h2&gt;

&lt;p&gt;The other half of this is Unity Catalog itself exposing an Iceberg REST catalog endpoint, so external engines can discover and query UC-governed tables using the standard Iceberg REST protocol — not just read the files, but resolve table metadata, snapshots, and schema through the same catalog interface Iceberg-native tools expect. That matters for governance: permissions, lineage, and audit logging stay centralized in Unity Catalog even when the consuming engine has never heard of Databricks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the seams still show
&lt;/h2&gt;

&lt;p&gt;This isn't a fully symmetric, zero-latency bridge, and it's worth setting expectations before promising a downstream team instant Iceberg access:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Metadata generation isn't instantaneous.&lt;/strong&gt; UniForm's Iceberg metadata is generated after the Delta commit, which means there's a small propagation window rather than true single-transaction atomicity across both formats. For most analytical workloads this is invisible; for anything expecting read-your-writes consistency from an Iceberg-side client, test it rather than assume it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Delta remains the primary write path.&lt;/strong&gt; The practical write pattern today is: write Delta, read Iceberg. Two-way write support directly through Iceberg clients into UC-managed tables is a newer and less battle-tested path than the read side — treat it as something to pilot, not something to build a critical pipeline on without validation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not every Delta feature has an Iceberg-side equivalent.&lt;/strong&gt; Deletion vectors and liquid clustering are genuinely Delta-native performance features; UniForm makes the table's data readable as Iceberg, it doesn't retroactively give the Iceberg spec those same optimizations. An Iceberg engine reading the table gets correct data, not necessarily every performance characteristic your Delta-side jobs enjoy.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The actual guidance
&lt;/h2&gt;

&lt;p&gt;Don't treat "Delta vs Iceberg" as a platform-wide decision that has to be made once and defended forever. Write Delta on Databricks, because that's where the write-side performance features live and where your pipelines already run. Turn on UniForm for any table an external engine needs to read, and use Unity Catalog's Iceberg REST endpoint to expose it, rather than standing up a separate export pipeline that will inevitably drift. The format war only matters if you're locked into a single-format assumption somewhere in your stack — increasingly, on Databricks, you don't have to be.&lt;/p&gt;

&lt;p&gt;Zephico is a &lt;a href="https://zephico.com/partners/databricks" rel="noopener noreferrer"&gt;Databricks Consulting Partner&lt;/a&gt;, and our &lt;a href="https://zephico.com/services/data-engineering-databricks" rel="noopener noreferrer"&gt;Databricks-certified engineers&lt;/a&gt; set this up for clients who need one governed copy of their data serving both Databricks-native and external Iceberg consumers. If your organization is maintaining duplicate Delta and Iceberg pipelines today, &lt;a href="https://zephico.com/contact" rel="noopener noreferrer"&gt;talk to us&lt;/a&gt; about collapsing them into one.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://zephico.com/blog/unity-catalog-open-formats" rel="noopener noreferrer"&gt;Zephico blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>databricks</category>
      <category>dataengineering</category>
    </item>
    <item>
      <title>Databricks SQL vs a classic BI warehouse: what actually changes for your analysts</title>
      <dc:creator>Zephico Technologies</dc:creator>
      <pubDate>Thu, 27 Aug 2026 22:45:41 +0000</pubDate>
      <link>https://dev.to/zephico/databricks-sql-vs-a-classic-bi-warehouse-what-actually-changes-for-your-analysts-79m</link>
      <guid>https://dev.to/zephico/databricks-sql-vs-a-classic-bi-warehouse-what-actually-changes-for-your-analysts-79m</guid>
      <description>&lt;p&gt;The pitch for Databricks SQL is usually made to platform teams: one copy of the data, no separate warehouse to sync, governance in one place. The question analysts actually ask is narrower and more practical — "I connect Tableau or Power BI to a warehouse today. What changes if that warehouse is DBSQL instead of Snowflake or Redshift?" Here's the honest answer, split into what changes and what doesn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  What doesn't change
&lt;/h2&gt;

&lt;p&gt;Analysts still write ANSI-standard SQL, and BI tools still connect the same way — ODBC/JDBC drivers, a host, a port, credentials. A dashboard built against DBSQL looks and behaves like a dashboard built against any other warehouse. Query results, joins, window functions, CTEs — none of that is exotic or Databricks-specific for the 95% of queries an analyst writes day to day. If the pitch to your BI team is "learn a new query language," that's the wrong pitch; there isn't one.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually changes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;There's no separate copy of the data to wait on.&lt;/strong&gt; A classic warehouse setup has ETL or reverse-ETL jobs landing data from source systems into the warehouse on some schedule — hourly, nightly, whatever the pipeline allows. DBSQL queries the same Delta tables your data engineers write to directly. When a pipeline finishes, the table is queryable immediately; there's no separate load step into a second system with its own lag. For analysts, this mostly shows up as fresher data and one less place to check when a number looks stale.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Compute is a warehouse you pick, not a cluster you configure.&lt;/strong&gt; SQL warehouses (serverless or classic) are sized T-shirt style — X-Small through 4X-Large — and scale out automatically under concurrent load, without an analyst or admin manually adding capacity. Query queuing under heavy concurrent use is real and does happen, but it's a warehouse-sizing conversation, not a "my dashboard is down" incident, and serverless warehouses start in seconds rather than the minutes a classic cluster used to take to spin up.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Photon changes what "fast" costs.&lt;/strong&gt; Databricks' vectorized query engine, Photon, is what makes SQL workloads on a lakehouse competitive with dedicated warehouse engines rather than noticeably slower — this matters because the historical knock on Spark-based SQL was interactive-query latency, and Photon is the specific answer to that complaint. It's on by default on SQL warehouses; there's nothing an analyst has to configure to get it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cost visibility moves to the query and the warehouse, not a flat subscription.&lt;/strong&gt; Classic warehouses often bill as a fixed-size commitment; DBSQL bills DBUs per warehouse per second it's running. That's usually good news for teams with lumpy usage, but it means an analyst who leaves a large warehouse running interactively all day is now a visible line item, not a rounding error inside an annual contract — cost governance (auto-stop timers, warehouse-level budgets) becomes a real setup task, not a nice-to-have.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where migrations actually snag
&lt;/h2&gt;

&lt;p&gt;The friction isn't SQL syntax — it's the small percentage of queries that lean on a source warehouse's proprietary functions, stored procedures, or scheduling/orchestration features baked into the old platform. Snowflake-specific functions, vendor-specific date arithmetic, or logic embedded in stored procs need translating, and that's real work, just not large work if it's scoped honestly up front. The other common snag is BI tool connectors: most major tools (Tableau, Power BI, Looker) have mature native DBSQL connectors today, but check any custom or embedded-analytics tooling before assuming parity.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual decision
&lt;/h2&gt;

&lt;p&gt;If the reason to move is "we already run our pipelines on Databricks and want to stop copying data into a second warehouse," DBSQL is usually a clean win — same SQL, fresher data, one governance model via Unity Catalog instead of two. If the reason is purely "warehouses are expensive, will this be cheaper," the answer depends entirely on your usage pattern, and it's worth modeling before committing either way.&lt;/p&gt;

&lt;p&gt;Zephico is a &lt;a href="https://zephico.com/partners/databricks" rel="noopener noreferrer"&gt;Databricks Consulting Partner&lt;/a&gt;, and our &lt;a href="https://zephico.com/services/data-engineering-databricks" rel="noopener noreferrer"&gt;Databricks-certified engineers&lt;/a&gt; run these evaluations for clients who are deciding whether to consolidate onto DBSQL or keep a separate warehouse. If you want a straight read on which side of that line you're on, &lt;a href="https://zephico.com/contact" rel="noopener noreferrer"&gt;talk to us&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://zephico.com/blog/databricks-sql-vs-bi-warehouse" rel="noopener noreferrer"&gt;Zephico blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>databricks</category>
      <category>dataengineering</category>
    </item>
    <item>
      <title>How much does custom software cost in 2026? An honest breakdown</title>
      <dc:creator>Zephico Technologies</dc:creator>
      <pubDate>Wed, 19 Aug 2026 06:23:02 +0000</pubDate>
      <link>https://dev.to/zephico/how-much-does-custom-software-cost-in-2026-an-honest-breakdown-1j32</link>
      <guid>https://dev.to/zephico/how-much-does-custom-software-cost-in-2026-an-honest-breakdown-1j32</guid>
      <description>&lt;p&gt;Ask five agencies what your app will cost and you'll get five refusals to answer. The honest reason: the range is wide and the drivers are mostly on the buyer's side. So instead of a fake number, here's how the number actually forms — and the levers that move it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The four things that set the price
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Scope, obviously — but really workflow count.&lt;/strong&gt; Software cost scales with the number of distinct workflows, not screens. "Customers order, ops fulfills, finance reconciles" is three products wearing one logo. The cheapest project decision you will ever make is cutting workflow two and three from version one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Integrations.&lt;/strong&gt; Every external system — payment provider, ERP, legacy database, that one SOAP API from 2009 — adds discovery, error handling and testing that dwarf the happy path. A form is cheap; a form that has to agree with NetSuite is not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Compliance and data sensitivity.&lt;/strong&gt; Health data, payments, or &lt;a href="https://zephico.com/services/gdpr-compliance-security" rel="noopener noreferrer"&gt;GDPR-regulated personal data&lt;/a&gt; add audit trails, access controls and review steps. Bolting this on later costs multiples of building it in.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Unknowns.&lt;/strong&gt; If nobody can describe what "done" looks like, you're paying for discovery either way — the only choice is whether you pay for it explicitly (cheap) or through rework (expensive).&lt;/p&gt;

&lt;h2&gt;
  
  
  Realistic ranges
&lt;/h2&gt;

&lt;p&gt;With senior engineers and honest scoping, as of 2026 we'd ballpark it like this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A focused MVP&lt;/strong&gt; — one core workflow, standard stack, no exotic integrations: &lt;strong&gt;low-to-mid tens of thousands of dollars&lt;/strong&gt;, a couple of months.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An internal platform or customer portal&lt;/strong&gt; — several workflows, real integrations, roles and permissions: &lt;strong&gt;high tens into low hundreds of thousands&lt;/strong&gt;, a quarter or two.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A product that is the business&lt;/strong&gt; — complex domain, years of iteration ahead: you're not buying a project, you're standing up a team; think in &lt;strong&gt;annual team cost&lt;/strong&gt;, not project cost.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Anyone quoting far below these bands is cutting something you'll pay for later — usually senior engineering time, which is exactly the ingredient that prevents rework.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why fixed-bid quotes backfire
&lt;/h2&gt;

&lt;p&gt;A fixed price on unfixed requirements just prices in the vendor's risk — you pay a premium &lt;em&gt;and&lt;/em&gt; create an incentive to fight scope changes. What works better: fix the budget and the team, keep scope adjustable, ship in slices you can evaluate. You keep the cost control fixed-bid pretends to offer, without the change-order theater.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to genuinely spend less
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cut scope, not seniority.&lt;/strong&gt; Three senior engineers outrun six juniors and produce less code to maintain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Choose boring technology.&lt;/strong&gt; &lt;a href="https://zephico.com/services/custom-software-development" rel="noopener noreferrer"&gt;Django, React, PostgreSQL&lt;/a&gt; — the boring stack is the one every future engineer can maintain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use timezone-aligned offshore engineering&lt;/strong&gt; — the cost lever that doesn't touch quality, if overlap hours are guaranteed rather than promised.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Buy instead of building&lt;/strong&gt; wherever the workflow isn't your differentiator. Sometimes the right answer to "build us an admin panel" is &lt;a href="https://zephico.com/blog/retool-internal-tools-when-to-use" rel="noopener noreferrer"&gt;Retool in three weeks&lt;/a&gt;, not a custom build in three months.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The most expensive software is the system you build twice. Whatever you spend, spend it on knowing what you're building — the rest of the invoice follows from that.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://zephico.com/blog/custom-software-development-cost-2026" rel="noopener noreferrer"&gt;Zephico blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>programming</category>
      <category>business</category>
    </item>
    <item>
      <title>Contract engineer vs. full-time hire: the real cost math</title>
      <dc:creator>Zephico Technologies</dc:creator>
      <pubDate>Wed, 12 Aug 2026 20:15:19 +0000</pubDate>
      <link>https://dev.to/zephico/contract-engineer-vs-full-time-hire-the-real-cost-math-4pdm</link>
      <guid>https://dev.to/zephico/contract-engineer-vs-full-time-hire-the-real-cost-math-4pdm</guid>
      <description>&lt;p&gt;When teams compare a contractor's monthly rate against a salary, the salary usually looks cheaper. That comparison is wrong on both sides, and since we &lt;a href="https://zephico.com/services/engineers-on-contract" rel="noopener noreferrer"&gt;place engineers on contract&lt;/a&gt; for a living, it's worth showing the honest version of the math — including the cases where hiring full-time is the right call.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a full-time senior engineer actually costs
&lt;/h2&gt;

&lt;p&gt;Take a senior engineer in the US or Western Europe with a headline salary somewhere around $150–200k. The real number is bigger:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Employment overhead.&lt;/strong&gt; Payroll taxes, health insurance, pension contributions, equipment, software seats, office or stipend. Finance teams typically model fully-loaded cost at 1.25–1.4× salary. Your $180k engineer costs the company roughly $230–250k a year before they write a line of code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Recruiting.&lt;/strong&gt; An agency fee runs 20–25% of first-year salary — $35–45k for one hire. Do it with an internal recruiter and you're paying their salary plus your engineers' interviewing hours, which are real hours that stopped producing software.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The vacancy itself.&lt;/strong&gt; Time-to-hire for senior engineers is commonly two to three months, plus notice period. That's a quarter of roadmap that either slips or lands on the rest of the team. Nobody books this to a budget line, which is why it's the most underestimated cost on the list.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The exit.&lt;/strong&gt; If the project ends, priorities shift, or the hire doesn't work out, you're into severance, notice periods and — in much of Europe — a legal process. The option to &lt;em&gt;stop paying&lt;/em&gt; has a price, and full-time employment doesn't include it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a contract engineer actually costs
&lt;/h2&gt;

&lt;p&gt;The monthly rate — and that's mostly it. No recruiting fee, no payroll overhead, no severance exposure. The engineer starts within days rather than months, and when the project ships you scale down without a process.&lt;/p&gt;

&lt;p&gt;There are real costs on this side too, and pretending otherwise would be salesmanship: onboarding time to your codebase, a management relationship to maintain, and knowledge that walks out when the contract ends unless you deliberately capture it. (We wrote about &lt;a href="https://zephico.com/blog/managing-contract-engineers-remote-teams" rel="noopener noreferrer"&gt;running contract engineers well&lt;/a&gt; — most of these costs are controllable.)&lt;/p&gt;

&lt;h2&gt;
  
  
  The break-even question
&lt;/h2&gt;

&lt;p&gt;The comparison isn't "rate vs. salary." It's:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Horizon.&lt;/strong&gt; If you're confident the role exists in three years, employment amortizes its fixed costs and wins. If the need is a project, a migration, a spike in roadmap — the fixed costs never amortize.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Certainty.&lt;/strong&gt; Hiring is a two-to-three-month commitment to a guess about your future roadmap. Contracting converts that guess into a monthly decision.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Knowledge type.&lt;/strong&gt; Core product architecture compounds in a long-tenured employee's head. Skills like "build the Databricks pipeline" or "stand up the Retool console" are transferable — you're buying the capability, not the tenure.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where we tell people to hire instead
&lt;/h2&gt;

&lt;p&gt;If the work is your core product and the horizon is years, hire. A staff engineer who has lived in your codebase for four years is worth more than any rotation of outsiders, and no honest staffing company should claim otherwise. Contract engineers win at the edges: projects with an end date, skills you need now but not forever, and teams that need senior capacity this month, not next quarter.&lt;/p&gt;

&lt;p&gt;Run the fully-loaded numbers for your own case before deciding. If the contract column wins, &lt;a href="https://zephico.com/services/engineers-on-contract" rel="noopener noreferrer"&gt;this is what our version looks like&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://zephico.com/blog/contract-engineer-vs-full-time-hire-cost" rel="noopener noreferrer"&gt;Zephico blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>hiring</category>
      <category>career</category>
    </item>
    <item>
      <title>Medallion architecture on Databricks: what actually matters</title>
      <dc:creator>Zephico Technologies</dc:creator>
      <pubDate>Wed, 12 Aug 2026 20:14:58 +0000</pubDate>
      <link>https://dev.to/zephico/medallion-architecture-on-databricks-what-actually-matters-2nom</link>
      <guid>https://dev.to/zephico/medallion-architecture-on-databricks-what-actually-matters-2nom</guid>
      <description>&lt;p&gt;Every Databricks pitch deck has the same three-layer diagram: bronze for raw data, silver for cleaned data, gold for business-ready tables. The diagram is fine. The problems start when teams treat it as an architecture instead of what it really is — a naming convention for a set of decisions you still have to make.&lt;/p&gt;

&lt;p&gt;Here's where those decisions actually bite, based on the lakehouse builds we've done.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bronze: append-only or you'll regret it
&lt;/h2&gt;

&lt;p&gt;The single most common mistake we see is teams "cleaning up" data on the way into bronze — deduplicating, fixing types, dropping malformed rows. It feels tidy. It's also how you lose the ability to reprocess history when your parsing logic turns out to be wrong.&lt;/p&gt;

&lt;p&gt;Bronze should be an append-only, schema-on-read record of exactly what the source system sent you, with ingestion metadata (&lt;code&gt;_ingested_at&lt;/code&gt;, &lt;code&gt;_source_file&lt;/code&gt;) and nothing else. Storage is cheap. The ability to replay six months of raw events after finding a bug in your silver logic is priceless.&lt;/p&gt;

&lt;h2&gt;
  
  
  Silver: this is where your real data model lives
&lt;/h2&gt;

&lt;p&gt;Silver is not "bronze but cleaner." It's where you commit to a data model: one row per entity, resolved keys, enforced schemas, quarantined bad records. Two rules that have served us well:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Every silver table gets expectations.&lt;/strong&gt; Delta Live Tables' &lt;code&gt;expect_or_drop&lt;/code&gt; (or plain constraint checks in a batch job) with quarantine tables for the failures. A silver table without declared expectations is a bronze table with a misleading name.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Silver models the domain, not the report.&lt;/strong&gt; If a table exists because one dashboard needs it, it belongs in gold. Silver tables should survive a BI-tool migration untouched.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Gold: let it be boring and duplicated
&lt;/h2&gt;

&lt;p&gt;Gold tables can be denormalized, redundant, and aggressively shaped for one consumer. That's the point. Resist the urge to build one "canonical" gold layer that serves every team — you'll end up with a committee-designed table that serves nobody and takes four teams to change.&lt;/p&gt;

&lt;h2&gt;
  
  
  The parts the diagram doesn't show
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Unity Catalog from day one.&lt;/strong&gt; &lt;a href="https://zephico.com/blog/unity-catalog-migration-what-it-takes" rel="noopener noreferrer"&gt;Retrofitting governance onto a running lakehouse is weeks of migration work&lt;/a&gt;; starting with it is an afternoon of setup.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost visibility before cost optimization.&lt;/strong&gt; Tag jobs by pipeline and team first. Most "Databricks is expensive" complaints turn out to be one forgotten streaming cluster or an oversized all-purpose cluster someone uses for ad-hoc SQL.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A reprocessing story.&lt;/strong&gt; Decide &lt;em&gt;before&lt;/em&gt; launch how you'll rebuild silver from bronze — full replays, partition-scoped replays, or versioned logic. If the answer is "we'd figure it out," you don't have a lakehouse, you have a pile of Parquet.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this is exotic. That's the real lesson: medallion architectures fail on discipline, not on technology. Get the boring parts right and the diagram takes care of itself.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://zephico.com/blog/medallion-architecture-databricks-practical-guide" rel="noopener noreferrer"&gt;Zephico blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>databricks</category>
      <category>dataengineering</category>
    </item>
  </channel>
</rss>
