<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Levent Çeliksan</title>
    <description>The latest articles on DEV Community by Levent Çeliksan (@levent_celiksan_617dcb736).</description>
    <link>https://dev.to/levent_celiksan_617dcb736</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4146071%2Fc33215b9-ba9c-4b43-b6fa-3b6b1bc16855.jpg</url>
      <title>DEV Community: Levent Çeliksan</title>
      <link>https://dev.to/levent_celiksan_617dcb736</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/levent_celiksan_617dcb736"/>
    <language>en</language>
    <item>
      <title>How I find edited copies of images across the web without crawling it</title>
      <dc:creator>Levent Çeliksan</dc:creator>
      <pubDate>Mon, 28 Sep 2026 04:44:05 +0000</pubDate>
      <link>https://dev.to/levent_celiksan_617dcb736/how-i-find-edited-copies-of-images-across-the-web-without-crawling-it-30ep</link>
      <guid>https://dev.to/levent_celiksan_617dcb736/how-i-find-edited-copies-of-images-across-the-web-without-crawling-it-30ep</guid>
      <description>&lt;p&gt;&lt;em&gt;Search engines have already crawled the visual web. Crawling it all again would duplicate that work, along with the hardware, energy and water behind it. I use their indexes as the first stage of a copy-detection pipeline and do all of the verification myself, on one Mac mini.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;If you publish photos or video, copies of your work end up on sites you have never heard of. They are rarely exact copies. They get cropped, resized, recompressed, color-graded, dropped into collages, re-titled in another language, or re-uploaded with the audio stripped out.&lt;/p&gt;

&lt;p&gt;Finding those copies is usually treated as a crawling problem. You fetch as much of the web as you can, fingerprint everything, and look your content up in that index. That is how the large systems work, and it is why they are expensive. You pay to crawl, store and re-crawl a large part of the internet before you can answer a single question.&lt;/p&gt;

&lt;p&gt;When I started building Sealify in 2023, I asked a different question: who has already crawled the web? Google, Yandex and Bing crawl and index billions of images, and they already offer reverse image search. What they do not offer is a precise answer. Their results mix true copies with pictures that only look similar.&lt;/p&gt;

&lt;p&gt;That gap shaped the whole design. I treat public search indexes as a broad but noisy first stage, and I put my own engineering into a strict, cheap second stage that runs locally. This post explains how it works, what it found in a real test, where it falls short, and why I think not crawling is also the responsible choice.&lt;/p&gt;

&lt;h2&gt;
  
  
  The web does not need another crawler
&lt;/h2&gt;

&lt;p&gt;When I designed Sealify, I asked a question engineers rarely ask: does this infrastructure need to exist at all? Not everything that is technically possible should be built.&lt;/p&gt;

&lt;p&gt;A copy-detection company that builds its own index has to crawl billions of pages, download the images and videos on them, fingerprint everything and store it. Then it has to repeat the cycle, because the web changes every minute. Google, Bing and Yandex already do all of this and keep their indexes fresh. A second index repeats their work. The same pages are fetched again, the same files are processed again, and the same data is stored again.&lt;/p&gt;

&lt;p&gt;A server is never just a line on a cloud bill. It is hardware that has to be mined, manufactured and shipped. It draws electricity every hour it runs, and data centers used about 1.5% of the world's electricity in 2024.[1] The building around it needs more energy, and often fresh water, to stay cool. In the United States alone, data centers consumed about 66 billion liters of water directly in 2023.[2]&lt;/p&gt;

&lt;p&gt;When the hardware is retired, it joins a stream of electronic waste that carries toxic materials such as lead and mercury. The world produced 62 million tonnes of e-waste in 2022, and less than a quarter of it was documented as properly collected and recycled.[3] Most of the rest is unaccounted for. Much of it is landfilled or handled by informal recyclers, where those materials can reach soil and water.&lt;/p&gt;

&lt;p&gt;In 2024, more than half of all web traffic was already automated.[4] One duplicate system looks small. If every new product rebuilt its own copy of the web, the waste would compound year after year, and the world would get nothing new in return. Reusing an index that already exists is like working a mine that is already open. Digging a second one beside it needs a very good reason, and copy detection does not have one.&lt;/p&gt;

&lt;p&gt;So Sealify never crawls the web. It asks the indexes that already exist, downloads only the few hundred thumbnails that matter for one query, and spends its own compute on the one job nobody else does for it: deciding which of those candidates is really a copy. I have not measured the savings in kilowatt-hours, and I will not invent a number. The principle holds either way. &lt;strong&gt;The cheapest and cleanest computation is the one you never run.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Retrieve with someone else's index, verify with your own
&lt;/h2&gt;

&lt;p&gt;I use the retrieve-then-verify pattern that is common in search systems. The difference is that I do not own the retrieval index. In copy detection, candidate generation is the expensive half, because it needs an index of the web. Verification is cheap by comparison. It compares one query against a few hundred candidates, and a single machine can do that in minutes.&lt;/p&gt;

&lt;p&gt;So the work splits like this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Breadth&lt;/strong&gt; comes from search engines that already index the web.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Precision&lt;/strong&gt; comes from a local verifier that I control, tune and measure.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Cost&lt;/strong&gt; stays close to a handful of API calls per scan plus local compute.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fatvzxnoqch3bzlm2kbw3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fatvzxnoqch3bzlm2kbw3.png" alt="Figure 1 · The pipeline. Amber stages depend on search engines; teal stages run locally. Nothing counts as a match until the local verifier accepts it." width="799" height="231"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Figure 1 · The pipeline. Amber stages depend on search engines; teal stages run locally. Nothing counts as a match until the local verifier accepts it.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Stage one: ask several engines at once
&lt;/h2&gt;

&lt;p&gt;For each query image, the scanner asks several engines in parallel: Google Lens (exact and visual matches), Google's classic reverse image search, Yandex and Bing. I reach them through a third-party search API that returns each engine's results as structured JSON, with automatic failover between providers.&lt;/p&gt;

&lt;p&gt;Asking several engines buys coverage. Each engine indexes a different slice of the web and ranks results in its own way. In the test run described below, Google Lens and Yandex together returned verified copies on 396 domains. Only 8 of those domains appeared in both. Either engine on its own would have missed roughly half of the copies.&lt;/p&gt;

&lt;p&gt;Every candidate is deduplicated, its thumbnail is downloaded, and it joins a pool for verification. At this point nothing is a match yet.&lt;/p&gt;
&lt;h2&gt;
  
  
  Stage two: decide locally what is really a copy
&lt;/h2&gt;

&lt;p&gt;A search engine answers the question &lt;em&gt;what looks like this?&lt;/em&gt; Copy detection has to answer &lt;em&gt;is this the same picture?&lt;/em&gt; Most false positives live in the gap between those two questions: another photo of the same actor, another poster from the same film, an image with a similar mood.&lt;/p&gt;

&lt;p&gt;My verifier scores every candidate twice, once on the scene and once on the face.&lt;/p&gt;
&lt;h3&gt;
  
  
  Scene score
&lt;/h3&gt;

&lt;p&gt;Several visual signals are combined into one weighted score. The dots show each signal's relative weight.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpjkfgp7twj4t5q9fmewn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpjkfgp7twj4t5q9fmewn.png" alt="Scene-score signals and their relative weights." width="799" height="314"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Scene-score signals and their relative weights.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The weighting surprised me. I expected the deep-learning embedding to do most of the work, and in practice it carries very little weight. A semantic embedding is built to say that two images show the same kind of thing. For copy detection that is the wrong property, because two different photos of the same person on the same red carpet look almost identical to it.&lt;/p&gt;

&lt;p&gt;Color distribution, palette and background behave better. They survive resizing, recompression and moderate crops, and they separate different shots of the same subject. Normalizing brightness before scoring makes them tolerant of filters and exposure changes as well.&lt;/p&gt;
&lt;h3&gt;
  
  
  Face gate
&lt;/h3&gt;

&lt;p&gt;When the query contains a face, a candidate must also pass an identity check using InsightFace with ArcFace embeddings. The acceptance rule is simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;accept if scene_score ≥ 0.65 and face_score ≥ 0.45
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The face gate removes the most common remaining error, a picture with the right colors and layout but a different person in it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Calibrated, with no new training
&lt;/h3&gt;

&lt;p&gt;I did not need to train a new model. On test cases I had checked by hand, the first version of the verifier made the right accept-or-reject decision about 60% of the time. Most of the improvement came from calibration: choosing which signals to use, tuning their weights, and setting the two thresholds. After calibration it made the right decision about 97% of the time.&lt;/p&gt;

&lt;p&gt;That figure is accuracy on a test set. The precision figure later in this post is a different measure, taken from accepted matches in real scans.&lt;/p&gt;

&lt;h2&gt;
  
  
  One real run, end to end
&lt;/h2&gt;

&lt;p&gt;Here is a single scan from July 2026. The query was the Turkish theatrical poster for &lt;em&gt;Beetlejuice Beetlejuice&lt;/em&gt;, released in Turkey as &lt;em&gt;Beterböcek Beterböcek&lt;/em&gt;. Every candidate in this run came through the API path described above.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqqv9h0r44u8wt66je6f0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqqv9h0r44u8wt66je6f0.png" alt="Figure 2 · Candidates vs. verified matches, by source (one scale: 761 candidates = full width)." width="800" height="273"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Figure 2 · Candidates vs. verified matches, by source (one scale: 761 candidates = full width).&lt;/em&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;The engines returned 761 candidates. The verifier accepted 643 of them, spread over 396 domains.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Google Lens exact matches were nearly all genuine (399 of 400). Yandex results were noisier, and 244 of 361 passed.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;The accepted matches had a median scene score of 0.96 and a median face score of 0.70.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The query carried Turkish title text. The accepted matches included the English-language poster with a different title block, tight crops with no text at all, merchandise listings, and a YouTube thumbnail in which the poster fills only the middle third of a three-poster collage. That thumbnail scored 0.66 on the scene and 0.84 on the face. The scene score alone was borderline, and the face gate confirmed it. With the scene threshold at 0.70, this copy would have been lost.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnxsaxlmklq1vtoy1ng7r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnxsaxlmklq1vtoy1ng7r.png" alt="Figure 3 · Where the verified copies were found (396 domains). The two engines barely overlap." width="800" height="102"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Figure 3 · Where the verified copies were found (396 domains). The two engines barely overlap.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Candidate generation took about 2 minutes. Verifying all 761 candidates took about 15 minutes. This was a large scan; the median across my test runs is 2.5 minutes. The split matters for scaling. The part I do not own is fast, and the slow part is local compute, which is cheap to add.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu583fqlcr5xi95h1op7h.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu583fqlcr5xi95h1op7h.png" alt="Figure 4 · Time spent in each stage (same run): ~2 min search, ~15 min local verification." width="800" height="102"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Figure 4 · Time spent in each stage (same run): ~2 min search, ~15 min local verification.&lt;/em&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The expensive half of copy detection is knowing which pairs of images to compare. The web's largest indexes already answer that.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Video uses the same pipeline
&lt;/h2&gt;

&lt;p&gt;A video becomes a set of frames, and each frame becomes a query image. The hard part is choosing the frames. Intros, credits and logos appear in many unrelated videos, and blank or blurry frames match nothing useful. The frame selector:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;skips the first and last 10% of the video, where intros, outros and logos usually sit;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;drops frames that are nearly one flat color, and frames that are too blurry (low Laplacian variance);&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;prefers frames with a visible face, because the face gate makes those matches reliable;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;searches about 10 frames, 5 at a time.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Because matching works on frames, a copy with the audio removed or replaced is still found. In one video test, the system found more than 1,050 verified matches on over 80 domains.&lt;/p&gt;

&lt;p&gt;The same engine also runs in the other direction. Before Sealify registers a new video, it checks whether that video is already online. Audio fingerprinting runs first, a Shazam-style recognizer and then AcoustID, because it is cheap and decisive when the soundtrack is intact. Frame search runs when it is not. If the video is already online, the registration is rejected.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running it all on one machine
&lt;/h2&gt;

&lt;p&gt;The whole platform runs on a single Mac mini with an M2 Pro chip, exposed to the internet through Cloudflare Tunnel. Its fixed cost is about $300 a year. Two design choices make that work:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;An admission-control queue.&lt;/strong&gt; A scan starts only when there is a free slot and enough memory and CPU headroom. Today that means 10 concurrent slots inside a 12 GB memory budget. Extra jobs wait their turn instead of taking the machine down.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Parallel verifier workers.&lt;/strong&gt; Five matcher processes run side by side. Each one claims its own port with a file lock, so a crashed worker restarts without disturbing the others. The models run on the Apple Silicon GPU.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Because the search stage runs on someone else's infrastructure, growth needs more verification compute. It never needs a crawler.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results across testing
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;3,000+&lt;/strong&gt; test runs&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;200,000+&lt;/strong&gt; candidate images scanned&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;10,000+&lt;/strong&gt; verified matches on 5,000+ domains&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;~98%&lt;/strong&gt; precision in manual review of the results&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The median scan took 2.5 minutes from upload to results. In manual review, about 2% of accepted matches turned out to be false positives.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it falls short
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Recall is capped by the engines.&lt;/strong&gt; If no engine returns a copy as a candidate, the verifier never sees it. Private groups, pages behind a login, and content the engines have not indexed are invisible.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Indexes lag.&lt;/strong&gt; A fresh upload may not show up in search results right away. Sealify covers this with a second layer: monitors that check chosen Instagram, TikTok and YouTube accounts every 15 minutes, and a monitor that watches known piracy sites. The index gives breadth and the monitors give freshness.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;It depends on third parties.&lt;/strong&gt; Engines change their results, prices and terms. Using several engines and keeping the verifier independent of the source limits the damage. It does not remove the risk.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Recall still needs a formal benchmark.&lt;/strong&gt; Precision comes from manual review, and recall has not been measured systematically yet. The next step is a seeded benchmark: publish controlled copies of test images (cropped, color-graded, mirrored, re-titled, re-encoded) at known URLs, then measure how many the system finds and how quickly.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What I actually built
&lt;/h2&gt;

&lt;p&gt;I did not need a new model. Every component here is known: reverse image search, perceptual hashing, color histograms, ArcFace, EfficientNet. What I built is the arrangement. Treat the web's largest indexes as a noisy first stage, verify strictly on your own hardware, and a problem that seems to need a crawler fits on a Mac mini. Nothing gets crawled twice.&lt;/p&gt;

&lt;p&gt;Sealify started as a way to defend creative work against piracy, and that is still its purpose. Behind every photo, video and design is someone's time, skill and effort. If technology makes that work easy to take, it should also make it easy to defend. Generative AI raises the stakes. People will soon need to know not only where their work was copied, but whether an image of them was altered or faked, and I expect the need for this kind of protection to grow.&lt;/p&gt;

&lt;p&gt;Two ideas sit under the whole project: respect for the people who create the work, and respect for the resources computing consumes. Good engineering spends those resources only where it creates something new.&lt;/p&gt;

&lt;p&gt;That is the kind of engineering I enjoy most. I look for infrastructure that already exists, built for some other purpose, and design the piece that makes it fit a new problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Patent Pending.&lt;/strong&gt; The system described here is covered by a US provisional patent application on content tracking and protection (63/923,341), filed in November 2025.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Next in this series: anchoring content fingerprints on Bitcoin, so every registration can be verified on a public block explorer with no fee for the user.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Sources
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;International Energy Agency, &lt;a href="https://www.iea.org/reports/energy-and-ai/executive-summary" rel="noopener noreferrer"&gt;Energy and AI&lt;/a&gt; (2025), executive summary.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Lawrence Berkeley National Laboratory, &lt;a href="https://eta-publications.lbl.gov/sites/default/files/2024-12/lbnl-2024-united-states-data-center-energy-usage-report_1.pdf" rel="noopener noreferrer"&gt;2024 United States Data Center Energy Usage Report&lt;/a&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;ITU and UNITAR, &lt;a href="https://ewastemonitor.info/the-global-e-waste-monitor-2024/" rel="noopener noreferrer"&gt;The Global E-waste Monitor 2024&lt;/a&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Imperva (Thales), &lt;a href="https://www.imperva.com/resources/resource-library/reports/2025-bad-bot-report/" rel="noopener noreferrer"&gt;2025 Bad Bot Report&lt;/a&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Levent Çeliksan&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Technical Founder &amp;amp; Senior Principal AI Engineer — Sealify, Inc.&lt;/p&gt;

&lt;p&gt;He has 19+ years of experience in backend systems. He designed and wrote the Sealify platform himself and launched it in June 2026. He is the inventor on six US provisional patent applications in content protection, deepfake detection and multimedia matching, and was invited to participate in IPHatch North America 2026.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://sealify.io" rel="noopener noreferrer"&gt;sealify.io&lt;/a&gt; · &lt;a href="https://linkedin.com/in/levent-celiksan" rel="noopener noreferrer"&gt;LinkedIn&lt;/a&gt; · &lt;a href="https://github.com/LeventCeliksan" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://leventceliksan.github.io/copy-detection/" rel="noopener noreferrer"&gt;leventceliksan.github.io/copy-detection&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>computervision</category>
      <category>python</category>
      <category>architecture</category>
    </item>
  </channel>
</rss>
