DEV Community

xiang li
xiang li

Posted on

How to Spot AI-Generated Images in 2026: A Forensic Multi-Signal Guide (Midjourney, FLUX.1 & DALL-E)

As diffusion architectures like FLUX.1, Midjourney v6, and DALL-E 3 achieve photorealistic fidelity, distinguishing synthetic visuals from authentic camera photographs has become a critical challenge for journalists, fact-checkers, designers, and web developers.

Many existing "AI detection APIs" operate as opaque black-box classifiers. They often cost thousands of dollars, suffer from high false-positive rates on compressed JPEGs, and require uploading sensitive user images to external third-party servers.

In this guide, we break down a deterministic, multi-signal forensic approach to inspecting digital images directly in the browser—analyzing metadata provenance, generative canvas geometry, and optical compression tells.


1. Signal One: Metadata & Provenance Traces (EXIF, XMP, C2PA)

Before running complex computer vision algorithms, examine the raw byte stream of the image file. Generative models and camera sensors leave distinct digital fingerprints.

A. Midjourney & Stable Diffusion Chunk Signatures

When downloaded directly or saved without heavy compression:

  • PNG tEXt Chunks: Stable Diffusion and ComfyUI write full generation parameters (Prompt, Negative Prompt, Seed, Sampler, CFG Scale) into the parameters metadata chunk.
  • Midjourney: Often embeds Description or software tags indicating the workflow version.
// Lightweight PNG text chunk parser in vanilla JS
function parsePngChunks(buffer) {
  const view = new DataView(buffer);
  let offset = 8; // Skip PNG header
  const chunks = {};

  while (offset < view.byteLength) {
    const length = view.getUint32(offset);
    const type = String.fromCharCode(
      view.getUint8(offset + 4),
      view.getUint8(offset + 5),
      view.getUint8(offset + 6),
      view.getUint8(offset + 7)
    );

    if (type === 'tEXt' || type === 'iTXt') {
      const chunkData = new Uint8Array(buffer, offset + 8, length);
      const text = new TextDecoder('utf-8').decode(chunkData);
      chunks[type] = text;
    }
    offset += 12 + length;
  }
  return chunks;
}
Enter fullscreen mode Exit fullscreen mode

B. C2PA Content Credentials & CAI

Platforms like OpenAI (DALL-E 3) and Adobe Firefly embed C2PA (Coalition for Content Provenance and Authenticity) metadata into the JPEG/PNG structure.

  • Look for JUMBF (JPEG Universal Metadata Box Format) boxes containing digital signatures (c2pa.actions.v2 stating c2pa.created).
  • If present, the image is mathematically certified as synthetic by the generator itself.

2. Signal Two: Generative Canvas Geometry & Resolution Fingerprints

Diffusion models are trained on specific latent bucket resolutions. While humans crop photos arbitrarily, synthetic images published across the web frequently retain default canvas dimensions:

Model Default Resolution Aspect Ratio
FLUX.1 [dev/schnell] 1024 × 1024, 832 × 1216 1:1, 2:3
Midjourney v6 default 1024 × 1024, 1456 × 816 1:1, 16:9
DALL-E 3 1024 × 1024, 1792 × 1024 1:1, 7:4
SDXL 1.0 1024 × 1024, 1152 × 896 1:1, 9:7

In contrast, real smartphone camera sensors shoot at standard hardware sensor ratios:

  • 4:3 native (e.g., iPhone 4032 × 3024, 48MP 8064 × 6048)
  • 3:2 native (DSLR full-frame sensors like Sony α7, Canon EOS)

If an uncropped image is exactly 1024x1024 or 1456x816 without camera EXIF, the probability of generative origin increases dramatically.


3. Signal Three: The Social Media Stripping Dilemma

A common failure mode for detection tools: Social media platforms (X / Twitter, Reddit, Instagram, Facebook) automatically re-encode uploads and strip all EXIF / C2PA metadata.

When metadata is stripped, naive tools give up or report "Inconclusive". A forensic analysis must check secondary visual tells:

  1. High-Frequency Over-Smoothing: FLUX and Midjourney often generate hyper-smooth skin transitions with synthetic micro-noise superimposed rather than true optical camera noise.
  2. Pupil & Specular Reflection Symmetry: Zoom in on eyes. Natural photos reflect the ambient lighting environment accurately across both eyes. Many diffusion models still hallucinate mismatched light sources or distorted pupil boundaries.
  3. Background Text & Glyph Coherence: Examine signage or license plates in the background. Synthetic text often mimics typographic form without spelling legible words.

4. Building a Privacy-First Verification Workflow

Uploading every image you come across to a remote server exposes private photos and introduces network latency.

To solve this, we built Check AI Free — a free, 100% client-side forensic inspection utility that evaluates metadata, canvas geometry, and visual tells directly in your browser:

  • Zero Server Uploads: Image bytes are analyzed locally via the Web File API.
  • Instant Multi-Signal Dossier: Breaks down prompt traces, EXIF presence, resolution matching, and social media compression status.
  • 1-Click Right-Click Inspection: We recently packaged this into the Check AI Free Chrome Extension, allowing you to right-click any image across X, Reddit, or news feeds and immediately view forensic signals without leaving your tab.

Summary Checklist

When inspecting an image online:

  1. ✅ Check file headers for PNG parameter chunks or C2PA JUMBF manifests.
  2. ✅ Compare canvas dimensions against known latent diffusion buckets (1024x1024, 1456x816).
  3. ✅ Look for missing camera sensor metadata (focal length, ISO, aperture) on high-res photos.
  4. ✅ Scrutinize optical coherence (specular reflections, text glyphs, ear/finger contours).

You can test any image today at Check AI Free.

What techniques or edge cases do you look for when debunking synthetic photos? Let's discuss in the comments below!

Top comments (0)