<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: VCh Tech</title>
    <description>The latest articles on DEV Community by VCh Tech (vch-tech).</description>
    <link>https://dev.to/vch-tech</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Forganization%2Fprofile_image%2F14982%2Fa19174dd-4196-4dce-8a96-5e929592d532.png</url>
      <title>DEV Community: VCh Tech</title>
      <link>https://dev.to/vch-tech</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/vch-tech"/>
    <language>en</language>
    <item>
      <title>Retry Is Not Recovery: A Practical Failure Model for Odoo Integrations</title>
      <dc:creator>Владислав Червяков</dc:creator>
      <pubDate>Tue, 29 Sep 2026 11:33:38 +0000</pubDate>
      <link>https://dev.to/vch-tech/retry-is-not-recovery-a-practical-failure-model-for-odoo-integrations-42d7</link>
      <guid>https://dev.to/vch-tech/retry-is-not-recovery-a-practical-failure-model-for-odoo-integrations-42d7</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1sjbqbod9vspb9ey1vx5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1sjbqbod9vspb9ey1vx5.png" alt=" " width="800" height="420"&gt;&lt;/a&gt;&lt;br&gt;
An Odoo integration job fails. The first response is often: put it back in the queue.&lt;/p&gt;

&lt;p&gt;That can repair a transient network failure. It can also create a second order, repeat a shipment update, or trigger an external side effect that already happened.&lt;/p&gt;

&lt;p&gt;The direct answer is simple: &lt;strong&gt;retry repeats an operation; recovery restores a known, consistent state.&lt;/strong&gt; Recovery may require a retry, but it may instead require replaying stored evidence, reconciling two systems, or stopping for a business decision.&lt;/p&gt;

&lt;p&gt;This distinction is especially important in Odoo integrations because a local transaction boundary does not describe what happened on the other side of an HTTP, REST, or XML-RPC call.&lt;/p&gt;

&lt;h2&gt;
  
  
  A queue is an execution mechanism, not a recovery model
&lt;/h2&gt;

&lt;p&gt;OCA's &lt;a href="https://github.com/OCA/queue/blob/18.0/queue_job/README.rst" rel="noopener noreferrer"&gt;queue_job documentation&lt;/a&gt; describes background jobs, isolated transactions, retry patterns, channels, and job dependencies. These are valuable building blocks. They let an integration postpone work, control concurrency, and retry a job after a retryable exception.&lt;/p&gt;

&lt;p&gt;But the queue cannot infer the business meaning of a failure.&lt;/p&gt;

&lt;p&gt;Suppose an Odoo job sends a create request to another system. The receiver commits the new record, but the response is lost. Odoo sees a timeout and marks the job as failed. Running it again is technically correct from the queue's point of view. From the business point of view, it may create a duplicate.&lt;/p&gt;

&lt;p&gt;Before choosing an action, classify the failure and establish what may already have happened.&lt;/p&gt;

&lt;h2&gt;
  
  
  A useful failure taxonomy
&lt;/h2&gt;

&lt;p&gt;I use the following taxonomy as an operational tool, not as a claim about Odoo's internal exception hierarchy.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Failure class&lt;/th&gt;
&lt;th&gt;Typical evidence&lt;/th&gt;
&lt;th&gt;Safe automatic action&lt;/th&gt;
&lt;th&gt;Main risk&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Transport&lt;/td&gt;
&lt;td&gt;timeout, connection reset, DNS failure&lt;/td&gt;
&lt;td&gt;retry only after checking side effects&lt;/td&gt;
&lt;td&gt;the remote system committed before the response was lost&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Authentication&lt;/td&gt;
&lt;td&gt;expired or rejected credentials&lt;/td&gt;
&lt;td&gt;stop, repair credentials, then resume deliberately&lt;/td&gt;
&lt;td&gt;repeated failures, lockout, or wrong identity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Schema mismatch&lt;/td&gt;
&lt;td&gt;missing field, wrong type, unexpected payload shape&lt;/td&gt;
&lt;td&gt;stop and fix the contract or mapping&lt;/td&gt;
&lt;td&gt;corrupt or incomplete data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;API contract mismatch&lt;/td&gt;
&lt;td&gt;model or method is absent, signature changed&lt;/td&gt;
&lt;td&gt;stop and verify the deployed API&lt;/td&gt;
&lt;td&gt;retrying a request that can never succeed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Source-data error&lt;/td&gt;
&lt;td&gt;missing required business data&lt;/td&gt;
&lt;td&gt;correct the source record, then replay&lt;/td&gt;
&lt;td&gt;bypassing validation and propagating bad data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Business conflict&lt;/td&gt;
&lt;td&gt;state transition is no longer allowed&lt;/td&gt;
&lt;td&gt;reconcile or route to a human&lt;/td&gt;
&lt;td&gt;overwriting a valid newer state&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Partial processing&lt;/td&gt;
&lt;td&gt;some steps committed, later step failed&lt;/td&gt;
&lt;td&gt;inspect checkpoints and reconcile&lt;/td&gt;
&lt;td&gt;duplicates or contradictory state&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The point of the table is not terminology. It is to prevent a single "retry" button from representing seven different operational decisions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Retry, replay, and reconciliation answer different questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Retry: can the same operation run again now?
&lt;/h3&gt;

&lt;p&gt;A retry repeats the same operation, usually with the same input, after a temporary problem. It is appropriate when the failure is plausibly transient and the operation is safe to repeat.&lt;/p&gt;

&lt;p&gt;"Safe" requires evidence. For a read, the answer is usually straightforward. For a write, ask:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Could the remote side already have committed?&lt;/li&gt;
&lt;li&gt;Can the previous result be identified?&lt;/li&gt;
&lt;li&gt;Does the receiver enforce a stable uniqueness or idempotency rule?&lt;/li&gt;
&lt;li&gt;Which downstream effects would run again?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Exponential backoff can reduce load. It cannot answer any of those questions.&lt;/p&gt;

&lt;h3&gt;
  
  
  Replay: can we reprocess the original evidence after a fix?
&lt;/h3&gt;

&lt;p&gt;Replay starts from a stored event, command, or source snapshot after correcting data, mapping, configuration, or code. It is not merely "run the failed function again."&lt;/p&gt;

&lt;p&gt;A controlled replay needs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the original input or an immutable reference to it;&lt;/li&gt;
&lt;li&gt;the contract version used at the time;&lt;/li&gt;
&lt;li&gt;a known restart point;&lt;/li&gt;
&lt;li&gt;a check for an existing result;&lt;/li&gt;
&lt;li&gt;an audit trail connecting the replay to the original attempt.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Replay is useful for deterministic faults. A renamed field or repaired mapping will not become correct through repeated retries, but the original event can often be processed successfully after the defect is fixed.&lt;/p&gt;

&lt;h3&gt;
  
  
  Reconciliation: do the systems agree now?
&lt;/h3&gt;

&lt;p&gt;Reconciliation compares actual state across systems without trusting the job history as the full truth.&lt;/p&gt;

&lt;p&gt;This is what finds an order that exists remotely while its local job says "failed," a shipment status that never returned, or a record created during partial processing. A reconciliation process normally matches stable identifiers, compares relevant fields or state, classifies differences, and either repairs them or creates an operator task.&lt;/p&gt;

&lt;p&gt;Retries handle execution attempts. Reconciliation handles reality.&lt;/p&gt;

&lt;h2&gt;
  
  
  External identifiers help, but they are not idempotency
&lt;/h2&gt;

&lt;p&gt;An external identifier creates a durable relationship between records in different systems. It lets an integration ask, "Have I already created the business object associated with this source record?"&lt;/p&gt;

&lt;p&gt;That is necessary for duplicate prevention, but it is not sufficient. Formal idempotency depends on where uniqueness is enforced, whether lookup and create are atomic, how concurrent requests behave, and which side effects occur outside the protected boundary.&lt;/p&gt;

&lt;p&gt;For example, this sequence is still racy:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Search the receiver for external ID X.&lt;/li&gt;
&lt;li&gt;Find nothing.&lt;/li&gt;
&lt;li&gt;Send a create request.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Two workers can pass step 2 before either creates the record. A database uniqueness constraint or receiver-side idempotency key closes a boundary that a preflight search alone does not.&lt;/p&gt;

&lt;p&gt;Use external identifiers to locate prior results and correlate logs. Do not describe them as a complete idempotency guarantee unless the full write path supports that claim.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two incident patterns from an Odoo-to-Odoo integration
&lt;/h2&gt;

&lt;p&gt;During my time at kt.team, I worked on the architecture and development of an integration between two Odoo 18 Community instances for SPL. One instance represented the marketplace side; the other handled supplier execution. Orders, purchases, shipments, stock data, messages, and attachments moved through separate modules and queued jobs.&lt;/p&gt;

&lt;p&gt;Two anonymized incidents illustrate why classification matters.&lt;/p&gt;

&lt;h3&gt;
  
  
  A field mismatch that retries could not repair
&lt;/h3&gt;

&lt;p&gt;A stock synchronization job expected &lt;code&gt;stock.move.line.qty_done&lt;/code&gt;, while the deployed contour used &lt;code&gt;quantity&lt;/code&gt; in that specific incident context. This is &lt;strong&gt;not a universal rule for every Odoo installation&lt;/strong&gt;; it was a mismatch between the integration's assumption and the actual model available in that contour.&lt;/p&gt;

&lt;p&gt;The job was deterministic. Every retry repeated the same invalid assumption. The correct response was to inspect the deployed model, fix the mapping, and then replay the affected work.&lt;/p&gt;

&lt;h3&gt;
  
  
  An API method that did not exist on the target
&lt;/h3&gt;

&lt;p&gt;Another job attempted to call a method on &lt;code&gt;product.product&lt;/code&gt; that was absent from the marketplace-side API. Connectivity and authentication were working. The contract was wrong.&lt;/p&gt;

&lt;p&gt;Again, retrying could not help. The integration needed to use the stock-update method actually exposed by the target, then reprocess the affected records under control.&lt;/p&gt;

&lt;p&gt;Neither incident was a "queue problem." The queue surfaced the failures and preserved work, but recovery depended on understanding the deployed contract.&lt;/p&gt;

&lt;h2&gt;
  
  
  A six-step recovery decision
&lt;/h2&gt;

&lt;p&gt;When an Odoo integration job fails, I use this sequence:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Preserve evidence.&lt;/strong&gt; Keep the job identifier, source record, correlation data, error class, and relevant request metadata. Do not expose credentials or personal payloads in logs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Classify the failure.&lt;/strong&gt; Separate transport, authentication, schema, API contract, source data, business conflict, and partial processing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Establish side effects.&lt;/strong&gt; Determine what may have committed locally and remotely before the failure became visible.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Locate an existing result.&lt;/strong&gt; Use external identifiers and receiver-side lookup before creating anything again.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Choose the operation.&lt;/strong&gt; Retry a transient safe attempt, replay preserved evidence after a fix, or reconcile actual state. Escalate business conflicts instead of overwriting them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verify the final state.&lt;/strong&gt; A green job is not enough. Confirm that both systems now agree on the business facts that matter.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The same model applies whether the transport is REST, XML-RPC, or another protocol. Protocol errors are observable symptoms; the recovery decision depends on business state and side effects.&lt;/p&gt;

&lt;h2&gt;
  
  
  Design for the failure path
&lt;/h2&gt;

&lt;p&gt;Recoverability is an architectural property. It comes from stable identifiers, explicit ownership, preserved inputs, observable checkpoints, constrained side effects, and a reconciliation path.&lt;/p&gt;

&lt;p&gt;Queues and retry policies remain useful. They are just one layer of the design.&lt;/p&gt;

&lt;p&gt;If the only recovery instruction is "retry the failed job," the system does not yet have a recovery model. It has an execution mechanism and an assumption.&lt;/p&gt;

&lt;h2&gt;
  
  
  Further reading
&lt;/h2&gt;

&lt;p&gt;The longer source article includes additional project context and an operational checklist: &lt;a href="https://www.vch-tech.com/blog/odoo-integration-error-recovery.html" rel="noopener noreferrer"&gt;How to recover an Odoo integration after failures&lt;/a&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;AI assistance disclosure:&lt;/strong&gt; I used ChatGPT to help structure and draft this article. I reviewed the technical claims against my published project material and the referenced OCA documentation, and I take responsibility for the final text.&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>odoo</category>
      <category>python</category>
      <category>architecture</category>
      <category>abotwrotethis</category>
    </item>
  </channel>
</rss>
