A flattened image does not contain a text node. It contains pixels that happen to look like letters. The poster, the menu photo, the label, the screenshot, the ad export: the type layer is gone, the style run is gone, and the only API you have is the raster. If you are choosing or building a feature that lets someone type a replacement and expects those pixels to become different letters without looking pasted on, you are not calling one model. You are running a stack, and the stack fails at its weakest stage.
The marketing sentence is always the same. Upload a picture, type a change, and the software erases the old words and draws new ones that blend in. That sentence hides a contract between stages. Detection can box the wrong strokes. The hole can be filled with a texture that does not belong. The style read can guess a family the renderer does not have. The new glyphs can be blurry, misaligned, or misspelled. Any one of those is enough to fail a human eye, which is the only judge that matters for a finished design.
This is the developer cut of that problem. I walk the stages the way you would sketch them before writing an interface: what each stage consumes, what it must emit, and how a bad emit poisons the next call. Then I write down the comparison harness, including the metrics I planned and did not publish. Then a matrix of the products that claim one or more of the stages. Then who should call which camp. I am not pasting a scoreboard. Where a metric was designed and not computed, I say the plan and withhold the number.
The practical gap is the one you already suspect if you have shipped a demo. A recorded demo picks a clean word on a flat color and stops the video before the proofread. A real job is a price, a date, a translated dish name, or a slogan on a photograph. The demo and the job are not the same function. Treat them as different tests or you will approve a pipeline that only works on the demo distribution.
The stack, stage by stage
Editing lettering in a flattened image is not one task. Treat the work as four stacked jobs — detect and OCR, inpaint, style match, render new text — because the composite is only as good as the weakest job. The data flow I want on a whiteboard is short enough to implement against:
input image → text detection and OCR → font and style analysis plus an edit mask → AI content fill, which is the inpaint → new text render with the style preserved → composite → edited image.
Most mainstream products stitch recognition, a fill, and a text update into that shape. The maturity of each module, and whether they share a mask and a string instead of each inventing their own, decides whether the edit survives contact with a person who can read. A stage that cannot name its input is not a stage. It is a prompt with a hope attached.
Here is the contract I would actually code between those jobs. None of these fields is optional if you want the run to be debuggable.
Box:
quad
raw_string
confidence
Style:
family_guess
size
color
stroke
shadow
warp
Mask:
hole
Edit:
target_string
operation
operation is either a replacement or a deletion. Deletion stops after the fill. Replacement continues into the renderer. If your product collapses all of that into one prompt string, you have not removed the jobs. You have made them impossible to unit test.
A second contract sits beside the schema: the source bytes are immutable. Write the edited raster to a new object. Keep the boxes, the mask, and the target string in a sidecar so a later pass can re-render without guessing what the model thought it saw. When a result looks wrong, you want to know which field was wrong. You cannot get that from a single flattened PNG and a shrug.
Detect and OCR
The first job finds the regions and reads them. Modern recognition combines deep learning, typically a convolutional front end with an LSTM sequence model, or a successor that does the same split of vision features and character order. On printed type with a clean ground it is solid. It is weak on low resolution, tiny fonts, handwriting, and busy backgrounds. That is not a footnote. Accurate detection is the foundation. Get the box or the string wrong and every downstream job fails in a way that looks like a rendering bug.
Think about what you store. An axis-aligned rectangle is the wrong primitive for a rotated sign or a word on a curved bottle. The rectangle includes background and misses glyph corners, so the mask is already dirty before the fill model runs. A quad, or a tighter polygon around the strokes, is the honest geometry. Multi-line grouping matters too. A headline and a subline that share a color are not one string, and a single box around both will ask the renderer to paint a paragraph into a title’s footprint.
Confidence is a gate, not a decoration. If the reader is unsure, the interface should show the raw string and let a person correct it before anyone inpaints. Silent failure here is expensive: the user “fixes” a word the model misread, the fill erases the real word, and the render dutifully draws the wrong correction in a beautiful font. Dedicated editors start by auto-boxing regions so a person can select one and edit it directly. That is the right product shape. Detection proposes. A human disposes. Fully automatic boxing with no confirm step is how you ship a tool that is confident and wrong.
Low resolution is a different failure from a tiny font. A tiny font at a high raster resolution may still have enough pixels per stem. A normal font photographed from across a street, then compressed, does not. Busy backgrounds break both the detector and the reader because edges of the artwork compete with edges of the strokes. Handwriting breaks the language model inside the reader, which was trained on type. None of these are reasons to skip the confirm step. They are reasons the confirm step exists.
There is also a product decision hiding in the box. Do you let the user edit only what was detected, or do you let them paint an extra region the detector missed? The second path is mandatory. Detectors miss reversed type, type in a shadow, and type that is the same color as a gradient behind it. A pipeline with no manual box is a pipeline that cannot recover. A pipeline that only has a manual box, and never proposes one, pushes the segmentation problem onto every user, which is the other failure mode. You want both: a proposal, and a brush that can add or subtract.
The string you show the user should be the string you will erase. Do not run a second, secret recognition pass later and erase a different set of glyphs. I have seen that class of bug in other document tools: the preview says one thing, the mask was built from another hypothesis, and a leftover stroke remains because the two hypotheses disagreed by a character. Pin the hypothesis. Persist it. Erase that.
Inpaint the hole
Once the mask exists, something has to invent the pixels that were behind the glyphs. Older inpainting such as PatchMatch copies nearby patches by nearest-neighbor search. It is fast, and it looks fine when the surround is stationary: a flat wall, a simple gradient, a repeating weave. It fails when the hole is large, when the texture is not stationary, or when a lighting falloff crosses the word. The copied patches tile, the lighting breaks, and you can still read the shape of the old word as a scar.
Those exemplar methods have largely been replaced by generative models, diffusion-based or GAN-based, that condition on the unmasked context and try to keep lighting and texture continuous. That is a real upgrade for photographs. It is also a new way to be wrong. A generative fill can invent a plausible brick pattern that was never there, or continue a highlight through the hole in a way that bends the light. For lettering removal, “plausible” is not the spec. The spec is “the surface that was already there, with the ink gone.”
Mask thickness is an engineering parameter, not a default you forget. Too tight, and a halo or a partial stroke remains. The next stage will sample that halo and treat it as part of the style. Too loose, and the fill eats a logo edge, a face, a rule line, or the neighboring word you meant to keep. Dilate and feather on purpose, and log the kernel you used. A removal demo will often describe this step as reading the surround and rebuilding whatever sat behind the glyphs. That description is the right mental model even when the implementation is a latent fill rather than a patch copy. Judge the result by whether a person can find the old baseline, not by whether the hole is “creative.”
Lighting gradients and texture are where this stage earns its keep or gives up. A word on a wrinkled cloth, a word on foliage, a word on a chrome edge: the conditional model has to continue a non-repeating signal through a hole shaped like letters. Letter-shaped holes are adversarial. They are thin, they have high-frequency boundaries, and they sit on top of the very signal you wish had been observed. Expect scars there. A flat printed sheet does not ask for this heroics, which is why so many demos choose it.
If you only delete, you are done after a successful fill. If you replace, the fill is an intermediate, not the product. It has to leave a surface the renderer can sit on. A fill that still contains ghost strokes will show through antialiased glyphs. A fill that changed the local noise profile will make the new word look like a sticker even when the font is right. Look at the hole before you look at the new word. Developers who only screenshot the final composite debug the wrong stage.
One more implementation note. Run the fill on the original pixels, not on a recompressed preview. Compression blocks along the mask boundary become part of the context the model trusts. If your UI shows a small proxy and your server fills the full raster, log both sizes in the sidecar or you will “fix” a bug that was only a proxy artifact. The same rule applies when the client downscales for upload and the server never sees the stems of the letters. You cannot inpaint detail you already threw away. Prefer a crop of the region plus a ring of context over a blind full-frame downscale, and keep the crop coordinates so the composite lands back on the original grid.
Match style before you draw
New lettering should match the original’s typeface, color, shadow, stroke, and distortion. That is its own stage, not a side effect of the fill. You are trying to separate a word’s content from its style, then apply the style to a string the user just typed. If you skip the separation, the renderer either uses a default font or asks a generative model to remember the font, and generative models are bad at remembering fonts.
Meta Research's TextStyleBrush extracts a text style — font, stroke weight, warp — from a single example and applies it to new content. That is one-shot style transfer. It is a research prototype, not a product in this test, but it shows the split is real: content in one vector, appearance in another. Commercial tools approximate the split with a built-in font match for a similar family, size, and color. Clipfly, for example, advertises editing lettering in the same font format. Approximation is the right word. A match against a library is not the same as a shader that reproduces a custom logo outline, a warped baseline, or a hand-painted terminal.
What the style stage must emit, if you want the renderer to be dumb and predictable:
- A family guess that maps to a font you can actually rasterize, plus a flag when the guess is weak.
- Size in pixels, measured on the quad, not on an axis-aligned box that inflated the height.
- Fill color and an optional stroke, sampled from the glyphs, not from the halo you forgot to dilate away.
- A shadow or extrusion description, even if the description is “none.”
- A warp or perspective mapped from the quad, so a word on a receding sign does not come back as a front-facing sticker.
If any of those fields is missing, do not pretend the render is style-preserving. Say the match failed and fall back to a visible default, or stop and ask. A silent fallback to a popular grotesque is how “AI edited my poster” becomes “AI replaced my brand.”
Length is part of style even though it feels like layout. The user’s new string is rarely the same width as the old one. The stage that owns style has to say whether the renderer may scale to fit, track tighter, wrap, or refuse. Scaling to fit changes the weight and breaks the match you just measured. Tracking tighter can collide with the decoration you did not mask. Refusal is a valid output. A pipeline that always squeezes is a pipeline that lies about fidelity.
Color sampling has its own trap. Glyphs are antialiased, so a naive average of the bounding box mixes ink with background and returns a muddy color. Sample the cores of the strokes, or the darkest or lightest decile depending on whether the type is positive or reversed. Reversed type, light letters on a dark photo, is where a sloppy sampler returns the background color and the new word vanishes. Store the polarity. The renderer needs it as much as it needs the family name.
Warped and logo typefaces are the honest limit of a library match. A match can find a near weight and a near width. It cannot become a custom wordmark. When the style stage’s weak-guess flag is set, the product should say so in the UI before the user spends a render. That flag is also what a later human proofread sorts on. You do not want those images in the same “looks done” pile as a clean reprint of a gothic on a white menu.
Render glyphs, do not hope the model spells
The last job paints the new content into the image so it belongs: layout, alignment, perspective, composite. There are two implementations, and they do not fail the same way.
The first asks a generative model to draw the letters. Asking a model to draw a specific string, the way DALL·E renders lettering, rarely produces clean, readable glyphs. You get extra strokes, fused characters, and spelling that drifted because the model does not have a character grid. It has a prior over pictures that contain something letter-shaped. That path can be useful for a texture or a sign that only has to look like language from across a room. It is the wrong path when a person will read a price.
The second path is the one that behaves like a word processor. Recognition produced a string. The user edits the string. A real text rasterizer draws glyphs from an actual font. The composite blends those glyphs through the warp you measured. In a purpose-built editor, clicking a region and typing a replacement makes the backend erase the old ink and imitate those letterforms with the new content. Imitation can mean “pick the nearest library font and match color,” which is honest, or “synthesize outlines,” which is research. Either way the letters are rendered, not dreamed. Spelling then fails only if the target string was wrong, which is a data bug you can test, not a sampling bug you cannot.
Composite order matters. Fill the hole first, on the original plate, then draw glyphs on top. If you draw glyphs and fill in one joint sample, you have gone back to the first path and you will not be able to correct a single character without resampling the whole patch. Joint sampling is why a one-word change alters the background noise and the neighboring letter you did not mean to touch. Keep the operations separable so a proofread can rerun the renderer without paying for another fill.
Perspective is a rasterizer feature, not a prompt adverb. If the source quad is a trapezoid, the glyph run has to be transformed into that quad after rasterization, or the rasterizer has to draw in that space directly. A rectangular overlay on a receding storefront is the sticker look, and no amount of color match hides it. Shadows and strokes belong in the same transform. A shadow that falls in screen space while the word falls in sign space is another sticker cue.
If you want one browser component to run every stage on a finished raster, that is the job people mean when they ask a tool to edit text in image. Judge that component as a pipeline with a confirmable string and a mask, not as a magic button. If it cannot show you the string it read, you cannot debug it. If it cannot leave the original file untouched, you should not run it on the only copy.
The render stage is also where critical information goes to die if you picked the generative drawer. A discount, a date, a phone number, a dosage: these are strings with a checksum in the reader’s head. A pretty wrong string is worse than a visible failure. Prefer a rasterizer that cannot misspell over a model that usually spells. When the style match is weak and you still choose the generative drawer, put a human proofread in the path and do not auto-publish.
The harness behind the comparison
What follows is a method, not a table of rates. The quantitative part of this project is a literature review plus a reproducible comparison plan. I leaned on official documentation and product notes, on academic papers, and on third-party reviews. Docs, papers, and reviews are restricted to the 2020–2026 source window, so this review year is the end of that window and the notes are not stale. The one-shot style prototype from the style stage is in that literature. So are the vendor descriptions of how each product wants to be driven.
The sample plan is 3 categories, about 20 images each, 60 total.
The categories are the ones that change which stage fails:
- Simple printed text. Black type on white, or any clean ink on a quiet ground. This is the distribution demos already live on. If a tool cannot do this, it cannot do anything harder.
- Text on complex backgrounds. Street signs, billboards, words with busy artwork or photography behind them. This is the inpaint stress test, and it is also a detection stress test when the artwork has its own edges.
- Special-style text. Handwriting, or commercial logo typefaces. This is the style-stage stress test. A library match can look confident and still be the wrong brand.
Each image uses moderately sized target lettering, not tiny lettering, so detection starts fair. I did not want a failure that was really “the stems were barely a few pixels tall and the task was unreadable to a person.” Tiny type is a separate experiment. Mixing it into this set would let a detector problem masquerade as a renderer problem.
The edit itself is fixed per image. Either replace the selected lettering with a new phrase, or delete the lettering. Use each tool’s intended method: recognition and a click, a direct replacement, or a prompt. Do not force a prompt-only tool through a click-to-edit UI it does not have, and do not refuse to use the auto boxes on a tool that spent its engineering budget on them. You are testing the product’s path, not your ability to misuse it.
These metrics were planned. This article does not report results for any of them.
- Recognition accuracy on the result. Run a standard reader on the edited region and check that the new string came out. From those checks, compute a correct-replacement rate. The rate is a method. I am not publishing one.
- Visual consistency. Human raters would score each edit from 1–5 on how well the new lettering blends in tone, lighting, and sharpness. An image-similarity metric such as SSIM could compare the original background with the repaired background. No average of that ordinal score appears here, and no similarity figure appears here. If you compute similarity, compute it on the repaired region plus a ring of context. A full-frame score stays flattering because most pixels never changed, which hides a bad local repair.
- Subjective aesthetics. Reviewers compare before and after and judge naturalness and visible flaws. An A/B blind pass is optional, not something I am claiming to have tallied.
- Automatic quality around the scar. Where you can, quantify the repair footprint with color balance or texture consistency in a band around the edit. Also planned. Also not a number in this piece.
The statistics plan matches the metrics. Per tool, compute the average replacement success and the average visual score, plus a variance so a single lucky image cannot define the tool. Where the sample allows, a t-test or a non-parametric test would check whether a difference between camps, or between a dedicated editor and a general scene fill, is more than noise. This article has no test statistic and no p-value. Without measured data, the expectation is stated in words. Vendor claims and industry feedback point the same way: font matching in the dedicated camp is expected to land high, and a general scene fill is expected to drift when the background is complex. Read that as an expectation, not as a rate I measured and then rounded away.
Reproducible steps are part of the method, or the comparison is a vibe. Log every tool’s real workflow. For an OCR editor: upload, automatic detection, click the region, type the replacement, download. For a desktop fill: record the selection geometry and the fill parameters. For a chat editor: save the prompt and the steps that indicated the region. Capture before and after for every scenario. Have more than one rater repeat the scoring so one person’s taste is not the metric. Hold the environment as still as you can: similar network conditions, the same image resolution across tools, the same target phrase on a given source. With more resources I would add type style and multiple languages as factors, and I would measure editing time and the feel of the UI. This write-up does not report those times.
A note on leakage, because this is where image comparisons usually lie to themselves. Do not tune the target phrase after you have seen which tool fails it. Do not drop an image from the set because one tool looked bad, unless you drop it for every tool and write down why. Do not score a tool on a screenshot of a screenshot if the other tools got the original file. The harness is boring on purpose. Boring is what makes a later number mean something. I am not inventing that later number in this text.
The same comparison, told as a full pass over the products rather than as a stage contract, is the piece on whether current models can edit text inside an image well enough to trust. This page stays with the pipeline, the harness, and the matrix.
The matrix, row by row
The dimensions that matter when you integrate one of these products, or when you refuse to, are narrower than a feature grid. What does it actually edit: a string, a region of pixels, or a background with no concept of glyphs? What is it good at when the stage matches the product? Where does it structurally give up? What is the price shape, without pretending a price is a quality score? Where do the bytes go? Who is the best caller?
Read the rows in that order. A dedicated-camp fit means a font-matching editor. A general-camp fit means a scene tool whose lettering is the unreliable stage. A removal-camp fit means delete only. One row never leaves the device and is not a lettering editor. One row only cuts out a background. One row is an unreleased demo for garbled glyphs in generated pictures. The pattern, once you sort that way, is stable. Some products are creative processors with no dedicated lettering editor. Some are friendly brush or prompt entry points. One family only deletes. The purpose-built lettering editors are the rows that prioritize font matching and automation.
| Tool | What it edits | Strengths | Limits | Price | Privacy | Best fit |
|---|---|---|---|---|---|---|
| Photoshop | Delete or replace selected elements, including lettering, and generate image content to fill the erased area. The fill step is Generative Fill. | Powerful scene context, fine control over selections and parameters, and it works with layer effects. The generative model family here is Firefly. | No dedicated text recognition or font matching. Generated lettering is often generic and not guaranteed legible. Steep learning curve, and the work is manual. | Subscription. | Edits upload to the Adobe cloud. Login is required. Copyright rules apply. | General camp, for professional scene work: complex scenes, object replacement, background extension. Not for precise lettering replacement; that still needs a manual font match on a real type layer. |
| Canva Express | Brush-select plus a text prompt to replace a selected element, for example deleting or replacing lettering. Magic Erase removes objects or lettering. Magic Edit is that prompt brush. | Simple and browser-based, with no install. It integrates with templates. Element replacement is fast on ordinary images. | Still in beta, with limited precision. Weak on complex backgrounds and fine detail. No auto-OCR, so you mark the region yourself. | Basically free. Advanced features need a paid Pro plan. A chat plugin can call it without a separate charge. | Online processing, and an account is required. The privacy policy stresses asset safety, but processing is still in the cloud. Output is subject to copyright and community rules. | General camp. Content creators doing quick element edits, or deleting and replacing lettering on flat design and social graphics, where pixel-perfect fidelity is not required. |
| ChatGPT + image editing | Upload an image and describe edits in natural language, including adding or removing lettering. In plugin mode it can call a desktop editor or a design app, for example blurring a background or replacing lettering through those plugins. | Flexible and iterative, with multi-turn consistency. The Images 2.5 model focuses on precise editing. Templates such as Sketch, plus prompt sharing, are part of the workflow. | Limited recognition and rendering. Expect errors or garbled spelling when lettering is drawn by the model. No exact font match. | Free, and Plus, Pro, and Enterprise tiers. Image editing is a general capability. Desktop plugins are reachable through the chat. | OpenAI is privacy-respecting in its framing, but uploads may be used for training, with an opt-out. Plugins require login and follow their own rules. | General camp. Creative prototyping and fast iteration: natural-language edits, quick slogans, color changes. Not for fine typographic proofing. |
| Pixelmator Pro | General editing and repair, including machine-learning content fill and object removal, plus built-in enhancement and canvas tools. | A friendly interface, Apple machine-learning features, and a smooth feel. Multi-layer and vector tools for creative work. | No dedicated text recognition or replacement. Mac and iPad only. | One-time purchase or a Mac App Store subscription. Cheaper than a full professional desktop subscription. | Mostly on-device. Whether anything is uploaded depends on what you do. No extra privacy exposure from the editor itself. | Routine editing for Mac designers: object removal and color adjustment. Not a lettering editor. |
| Remove.bg | Automatic background removal only. No lettering edit. | Focused and efficient. Fast, accurate deep-learning segmentation. | Single-purpose, and unrelated to lettering edits. | Basic free. Paid plans cover batch and high-resolution removal. | Web or API use requires an upload. The policy says content is not stored, or is kept only briefly in order to process. | Background removal for commerce and cut-outs. Useful as preprocessing before manual lettering work. Not a lettering editor. |
| Runway ML | Cloud image and video generation and editing, including object removal, expansion, and video inpainting. | A creation platform for stills and motion, with collaboration and an API. Frontier models include Gen-4.5 and GWM. | Aimed at creative generation and video, not dedicated lettering-object editing. Some advanced features need paid credits. | Free basic tier, pay-as-you-go, plus team and enterprise plans. | A cloud service that caches processed content. Check the terms on using assets for product improvement and on commercial licensing. | Creative video and interactive work. It can delete lettering as a side effect. It is not a font match. |
| ClipDrop | Cleanup removes a selected object, person, or lettering. Text Remover deletes lettering and, as the product describes it, rebuilds what was behind the text. | Simple, generally natural results, including on complex backgrounds, plus a free allowance. | Can only delete. It cannot replace or add lettering, cannot keep the original style, and cannot re-typeset. Needs a network. | Free daily quota. Premium is a subscription. | Account login for the web app. Images are uploaded to the cloud under Stability AI policies. | Removal camp. Clearing unwanted lettering, watermarks, or blemishes. Replacement is a different tool’s job. |
| Fotor | Auto-OCR on text regions. Click a detected region and type the replacement. The backend removes the old lettering and imitates the same letterforms with the new content. Prompt-based edits are also available. | Purpose-built for lettering in images. Auto-matches font, color, and effects. The flow is close to a word processor. It can delete, replace, and add lettering. | Needs a network, and server time varies. Weak on tiny type, handwriting, and warped type. | Free tier can watermark. Premium unlocks batch and high-res. | Images are transferred encrypted. The policy claims no retention for online editing. | Dedicated camp. Content marketing, ads, promo pages, social images, and typo fixes where font fidelity matters. |
| ReWords AI | Detects existing lettering, removes it, rebuilds the background, and renders a translation or replacement with matching font, color, and size. Built for finished images when there is no source file. | Aimed at flyers, posters, menus, and labels in the browser. The everyday job is localization and correction without rebuilding the layout. | Results depend on image quality, font complexity, and background texture. Complex photographic backgrounds are best-effort. Not for IDs, invoices, or official documents. | Free first edit after sign-in. Credit packs and subscriptions cover volume and higher-resolution exports. | Cloud processing. Read the privacy terms before a sensitive upload. | Dedicated camp. You do not have the design file, and you need the wording changed rather than the scene reinvented. |
| Textify (Storia.ai) | Mainly fixes garbled lettering inside generated images, replacing wrong glyphs with a meaningful string. | Designed for generated-art workflows. It improves readability somewhat and offers a one-click replacement. | Only pseudo-text in generated images, not real photos. Current results are limited, and examples still show spelling errors. | Not officially released. It is a project demo. | No public privacy policy. The project is still in development and testing. | Hobbyists repairing garbled lettering in generated pictures such as Stable Diffusion. Immature, and not for real photos. |
| Baidu Netdisk AI retouching | Adds image-text recognition and editing inside the drive app. Select a region and change font, color, and size. | Short steps. Auto-detects regions and offers edit controls. Advertises editing in the same font format. Saving and sharing stay inside the same ecosystem. | The app is required. Accuracy depends on scene complexity. Output quality is thinly verified in public. | Built into the drive app. Access may depend on a VIP tier. Described as currently free. | An internal feature. Processing may happen on the vendor’s servers under the drive’s general privacy rules, including encrypted storage. | Dedicated camp. Quick replacement for people already in that ecosystem, such as ads and layout notes. Watch the output. |
That finished-image row — detect, remove, rebuild, render, in a browser, with no design file — is the one you open when the export is all you have. The matrix is the integration view. The passes themselves were less tidy than any single row.
What the passes showed, without a scoreboard
Automatic lettering edits inside images are not a mature, single-purpose capability, and coverage is still split by what each product actually automated.
The dedicated camp is the one that preserves font style and turns a replacement around quickly. That matches the architecture: a box, a string, a fill, a rasterizer that imitates letterforms instead of sampling them. The general camp wins on image consistency and on creative flexibility, and it lacks fine control of type. A scene fill can remove lettering and rebuild the background without an obvious seam, and the letters it invents are still not style-consistent. A natural-language editor can add or change wording in one pass, and it still returns typos or warped glyphs often enough that you cannot skip a proofread. Those are qualitative splits. They are not recognition rates, not similarity scores, and not significance tests.
One observation held across the set, and it maps directly onto the inpaint stage. Lettering on complex backgrounds — texture, lighting gradients — failed or left artifacts far more often. Printed lettering on a quiet ground came out well in almost all of the tools. If your image is in the easy category, you are in the easy zone, and you should not brag about it as if it were the hard zone. If busy photography sits behind the words, expect a fight, and budget a human look at the hole before you accept the composite.
I am not going to dress that observation up as a computed footprint. The color-balance and texture checks in the harness are how I would quantify “artifact” on the next pass. They were not run to a published figure for this piece. The human ordinal score was not averaged. What I will stand behind is the direction: quiet grounds are kind to every camp, and textured grounds expose the fill. The style stage has its own direction, separate from the fill. Logo typefaces and handwriting are where imitation stops being imitation and becomes a different word in a nearby font. If the brand lives in the outline, do the change on a real type layer or do not do it.
Spelling failures deserve a developer’s reading, not a shrug. When a rasterizer misspells, the target string was wrong or the font did not contain the character, and you can assert on that in a test. When a generative drawer misspells, the sample drifted, and the next sample may drift differently. The second bug is not something you patch with a prompt adjective. You route critical strings away from that drawer. The passes made the routing obvious even without a rate beside it.
The short answer the stack forces
Shipped products fall into three camps, and each camp wins on a different axis.
Dedicated editors lead on font fidelity and speed. They are built to keep the original style, to box the words, and to let you type. That is the path that respects the stage contract: a string you can see, a mask, a fill, a rasterizer.
General generative tools lead on scene context and on creative changes to the picture as a picture. They understand more of the surrounding scene than a font matcher does. Their generated lettering is often blurry, misaligned, or misspelled, because lettering is the stage they refuse to isolate. Use them when the job is the scene. Do not use them as a type engine.
Removal-only tools delete lettering cleanly and cannot put words back. They are a fill stage with a UI, not an editor. That is useful, and it is not the feature the headline promised.
No single tool does everything well. The headline feature everyone wants — type new text and have it look native — is still the weakest link. In the passes I cared about, clean removal was easy. Convincing replacement was not. The stack explains why. Removal can stop after a decent fill. Replacement has to survive detection, style, and a renderer, and readers notice a bad glyph faster than they notice a slightly wrong brick.
Do not read the camp split as “only the dedicated camp is allowed.” Scene reconstruction, object removal, and extending a background are real jobs, and the general camp is built for them. The mistake is asking a scene model to also be a type setter, or asking a type setter to invent a new background object. Pick the camp that owns the stage you are actually calling. If you need two stages from two camps, run them as two calls and keep the intermediate, so you can throw away the bad call without rerunning the good one.
Who should run which camp
If you are a content creator, the usual job is small and specific. A discount changes. A date changes. A label gets a translated line. Prefer a dedicated editor for that. The flow feels like a word processor: select the words, type, let the backend rebuild the ground and keep the style. The barrier is low because you are not asked to become a scene compositor. For a price, a date, or a translated line on an export, edit text in image on the finished file instead of reconstructing the layout from scratch. For heavier creative changes — overall color, background objects, a different composition — add a general scene edit, and plan to touch up the output. Avoid relying on directly generated in-image lettering to carry critical information. The error mode is wrong characters that still look intentional. That is worse than an obvious hole.
If you are a designer, the bar is precise visual consistency, and a font match you do not control is not precise. Use a scene fill as an aid: remove the old lettering with it, then set the new wording yourself on a type layer so the family, the kerning, and the brand rules are yours. A chat product’s desktop-editor extension can streamline part of that flow, especially when you are still deciding what the new line should say. A browser brush-and-prompt edit can speed prototypes and internal reviews. Check the pixels before they leave the review. Treat automatic replacement as an assistant, not as a one-stop type department. The assistant is allowed to be wrong. Your file is not.
If you are on a legal or compliance team, the risk is privacy and copyright, and it is a data-flow risk rather than a style opinion. Every cloud service in this set — chat, browser design tools, dedicated web editors — requires the image to be uploaded. Upload implies a leak surface and, on some of those services, a training-use question. Read the vendor’s privacy terms. Confirm that retention, region, and authorization match the policy you already apply to customer artwork. An opt-out that exists only in a personal account setting is not an enterprise control until you have tested that it binds the workspace you actually use. Output raises a second issue: a rewritten logo outline can step on the original creator’s rights, and a generated sentence can state a fact you did not approve, including a price or a claim. For official documents and for ID photos, do not use an automated tool to edit the lettering. The downside is a dispute, not an ugly kerning pair.
On-device editing is the exception that proves the rule. If the pixels never leave the machine, the training-use question changes shape. The moment a “local” app calls a cloud fill for the hard stage, you are back in the upload regime for that stage. Ask which stage is local. Do not accept the word local for the whole pipeline because the chrome is a desktop app.
The safest approach is a mixed toolkit, with the original file kept untouched. For a simple string change, start with recognition plus a dedicated editor. For a creative edit, or when the source file is missing and the scene itself has to change, use a desktop fill or a chat model as a first pass, then proofread. For clearing unwanted lettering, use a removal tool, and do not expect it to type. Stay critical of the output. Compliance teams should write a policy for image editing and for the review of generated content, including who is allowed to upload customer files and which document classes are out of scope no matter how good the demo looked.
A practical policy is short enough to fit on one page. Keep the source object immutable. Store the edited object beside it, not instead of it. Require a human proofread when the string is a price, a date, a name, a legal line, or anything regulated. Forbid automated edits on identity documents and on official records. Record which vendor processed the file, because the privacy review is per vendor, not per category. None of that requires a new model. It requires the sidecar the stack section already asked for: operation, target string, and whether a person confirmed the read.
What I would actually ship
The models can change pixels that used to be letters. Whether they can do it well is a separate question, and the answer still depends on which stage you needed. Dedicated editors, built to detect, erase, match, and rasterize, are the ones that take the headline feature seriously. General-purpose generators still cannot reliably draw lettering that belongs in a design. Removal tools remain good at the job their name implies, which is the fill, and only the fill.
If I were reviewing a pull request that added “edit the text in the picture” to a product, I would block it unless the interface showed the detected string, the mask, and a way to edit the string before the fill ran. I would block a path that sent a price or a legal line through a generative drawer. I would block any path that overwrote the only copy. I would ask which camp owns each stage, and I would not accept one vendor checkbox as an answer for all of the stages. I would also ask what the harness is. A single hero screenshot is not a harness. The category split in the method section is the minimum: quiet type, type on a hard ground, and type whose style is the brand. If those situations were not tried, the feature is a demo.
Ship the mixed toolkit. Keep the original. Proofread the glyphs at actual pixel size, not in a thumbnail. Write the policy before someone pastes an identity document into a cloud fill because the button was convenient. The stack is understandable. The weakness is specific. It sits in the last mile, where new letters have to look as if they had been printed with the old ones, on a surface the fill had to imagine. That mile is still the hard part, and any roadmap that marks it done because a clean word on a white card looked fine has tested the easy category and called it the product.
Top comments (0)