<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Carlosjv</title>
    <description>The latest articles on DEV Community by Carlosjv (@carlosjv91).</description>
    <link>https://dev.to/carlosjv91</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4166742%2F34febb70-42af-4b46-ae3d-e1d71ad09694.png</url>
      <title>DEV Community: Carlosjv</title>
      <link>https://dev.to/carlosjv91</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/carlosjv91"/>
    <language>en</language>
    <item>
      <title>GroundLens: a tiny local model with a button to leave the screen</title>
      <dc:creator>Carlosjv</dc:creator>
      <pubDate>Tue, 06 Oct 2026 15:48:59 +0000</pubDate>
      <link>https://dev.to/carlosjv91/groundlens-a-tiny-local-model-with-a-button-to-leave-the-screen-5708</link>
      <guid>https://dev.to/carlosjv91/groundlens-a-tiny-local-model-with-a-button-to-leave-the-screen-5708</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-week1-2026-10-05"&gt;Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass&lt;/a&gt;.&lt;/em&gt; &lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;The most useful button in GroundLens is &lt;strong&gt;“Take this mission outside.”&lt;/strong&gt; It gives the screen a stopping point.&lt;/p&gt;

&lt;p&gt;GroundLens learns a small color palette from an outdoor photo, lets you choose one color, and asks you to find it in two different places. A leaf and a stone. Bark in sunlight and bark in shadow. The task is to notice what a label would miss: light, texture, distance.&lt;/p&gt;

&lt;p&gt;The intended audience is anyone who would enjoy a five-minute observation pause, including people who do not know the names of the plants around them. A garden, a courtyard or a nearby accessible outdoor spot is enough. The app does not require a long hike, species expertise, a camera permission or a GPS fix.&lt;/p&gt;

&lt;p&gt;The interaction has three parts:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Notice:&lt;/strong&gt; choose a photo and inspect its locally learned colors.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Leave:&lt;/strong&gt; take one observation prompt outside and put the screen away.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep:&lt;/strong&gt; return with a sentence, optionally saved only on your device.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;There are no streaks, points or notifications. Five minutes is a suggestion, not an achievement threshold. That is a design decision; I have not measured whether it reduces screen time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://groundlens.zrm7tv2zmj.chatgpt.site" rel="noopener noreferrer"&gt;Open GroundLens&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For a quick walkthrough, use the clearly labeled &lt;strong&gt;synthetic garden&lt;/strong&gt; already in the app. Choose a palette swatch, select &lt;strong&gt;Show learned regions&lt;/strong&gt;, then &lt;strong&gt;Take this mission outside&lt;/strong&gt;. Return with &lt;strong&gt;I'm back — capture one detail&lt;/strong&gt; to try the local notebook. The initial illustration is a reproducible demo, not a photograph or a claimed field trial.&lt;/p&gt;

&lt;p&gt;For your own scene, choose a JPG, PNG, WebP or AVIF image under 12 MB. The browser decodes it locally. The application does not upload the image or request your location.&lt;/p&gt;

&lt;p&gt;The method panel includes &lt;strong&gt;Download the complete offline app + source&lt;/strong&gt;. Extract it and open &lt;code&gt;dist/index.html&lt;/code&gt;; there is no installation or inference service to configure. Local-storage behavior for files varies between browsers, so the notebook also offers JSON export.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/carlosjunquerovila-afk/groundlens" rel="noopener noreferrer"&gt;GroundLens repository&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Apache-2.0. The project and repository were started on October 6, 2026 for this challenge. The learning engine is &lt;code&gt;dist/model.js&lt;/code&gt;, the interface is &lt;code&gt;dist/app.js&lt;/code&gt;, and numerical checks live in &lt;code&gt;tests/model.test.cjs&lt;/code&gt;. There are no runtime packages or pretrained weights to download.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;h3&gt;
  
  
  A model small enough to inspect
&lt;/h3&gt;

&lt;p&gt;The core is &lt;strong&gt;classical unsupervised machine learning: k-means++ color clustering&lt;/strong&gt;, implemented in open JavaScript. It is not a foundation model or a species recognizer. The model learns anew from each photograph; the learned centers determine both the palette and the optional segmented view.&lt;/p&gt;

&lt;p&gt;The pipeline resizes the image to at most 720 pixels on its longest side and samples at most 6,500 pixels. RGB values are converted to CIELAB, where Euclidean distance is more useful for this color-reconstruction task than raw RGB distance.&lt;/p&gt;

&lt;p&gt;Every fifth sampled pixel is held out. Three deterministic initializations—seeds 26, 91 and 2026—fit up to five clusters to the remaining pixels. The selected restart has the smallest training objective. Each restart has a maximum of 35 Lloyd iterations.&lt;/p&gt;

&lt;p&gt;The app then predicts the held-out pixels by nearest center. It displays their mean ΔE76 error beside a simple baseline: reconstructing every pixel with the training mean color. This exposes what the model actually accomplished rather than attaching an unexplained confidence percentage to a photograph.&lt;/p&gt;

&lt;p&gt;The observation sentence is a fixed template populated with the selected learned color. A small hand-written vocabulary supplies an approximate color name alongside its hex value. No language model invents a description of the scene.&lt;/p&gt;

&lt;h3&gt;
  
  
  What the tests establish
&lt;/h3&gt;

&lt;p&gt;The same JavaScript model runs in the browser and in Node. Reproduce the numerical checks with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;node tests/model.test.cjs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Seven tests passed on Node v24.19.0. They cover RGB/Lab round trips, a known two-color mixture, a constant image, invalid inputs, deterministic optimization, exact split/proportion accounting and 30 seeded synthetic color mixtures.&lt;/p&gt;

&lt;p&gt;Across those 30 mixtures, mean held-out reconstruction error was &lt;strong&gt;3.97 ΔE76&lt;/strong&gt;, compared with &lt;strong&gt;41.88&lt;/strong&gt; for the single-mean baseline. All 30 cases improved on that deliberately simple baseline. Per-case results are included in &lt;code&gt;tests/results.json&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That result has a narrow meaning: the implementation can recover clustered synthetic colors. It is &lt;strong&gt;not&lt;/strong&gt; evidence of species accuracy, ecological insight, generalization to future photos or improved wellbeing. Nearby pixels are correlated, so even the within-image holdout shown in the app is a diagnostic, not an independent field-validation set.&lt;/p&gt;

&lt;h3&gt;
  
  
  Boundaries I kept visible
&lt;/h3&gt;

&lt;p&gt;A nearly uniform picture triggers a low-variation message rather than a richer invented interpretation. Five clusters can miss small details. Lighting, exposure and camera processing change the palette. The segmented image has color regions, not identified objects.&lt;/p&gt;

&lt;p&gt;No outdoor trial or user study has been performed. Browser interaction and WebMCP validation were unavailable in the static preview environment; model tests and JavaScript syntax checks passed. Those are remaining validation gaps, not tests silently counted as successful.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;For this project, openness is practical: the complete learning algorithm is short enough to read, change and run without asking a provider for access.&lt;/p&gt;

&lt;p&gt;A teacher could replace the mission wording. A developer could compare sampling strategies or change the number of clusters. A privacy-conscious user can inspect the application and run the downloaded folder offline. The model has no credentials, quota or remote inference dependency.&lt;/p&gt;

&lt;p&gt;Photos remain in memory. Only an explicitly saved text note, selected color and timing information enter device storage. The application has no analytics or network requests of its own; the hosted page still requires ordinary requests to its hosting provider. Downloading the bundle removes that hosting dependency for subsequent offline use.&lt;/p&gt;

&lt;p&gt;The open implementation matters because learning colors locally is the whole interaction. If it disappears behind a paid inference endpoint, the app becomes less useful precisely where it is supposed to work: away from a desk, without having to send a private photograph anywhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  Development Disclosure
&lt;/h2&gt;

&lt;p&gt;An autonomous AI coding agent generated the implementation, tests, documentation and this write-up at my request. This is disclosed as &lt;strong&gt;Fully Autonomous&lt;/strong&gt;. No human editing or human-run field experiment is claimed.&lt;/p&gt;

&lt;p&gt;The algorithm follows k-means/Lloyd's iterative quantization and Arthur–Vassilvitskii k-means++ initialization, with sRGB/D65 CIELAB conversion. The implementation is newly written; no third-party source code or pretrained weights are bundled.&lt;/p&gt;

&lt;p&gt;The entry is for the overall challenge. No partner-technology prize is claimed.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>hf26challenge</category>
    </item>
    <item>
      <title>Receipt Before Claim: a 4% score that did not mean 96% wrong reasoning</title>
      <dc:creator>Carlosjv</dc:creator>
      <pubDate>Tue, 06 Oct 2026 14:54:45 +0000</pubDate>
      <link>https://dev.to/carlosjv91/receipt-before-claim-a-4-score-that-did-not-mean-96-wrong-reasoning-517o</link>
      <guid>https://dev.to/carlosjv91/receipt-before-claim-a-4-score-that-did-not-mean-96-wrong-reasoning-517o</guid>
      <description>&lt;p&gt;This is a submission for the Kaggle Benchmarking Challenge:&lt;br&gt;
&lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;https://dev.to/challenges/kaggle-2026-09-23&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;A successful build is not a live deployment. A receipt is not an acceptance. A result on a development set is not an external-test result. These distinctions matter whenever a model reports whether work is finished.&lt;/p&gt;

&lt;p&gt;I built Receipt Before Claim around four such boundaries: deployment status, external evaluation, application decisions, and version-specific test results. Each boundary has three synthetic evidence packets: one supports the claim, one contradicts it, and one leaves it unknown. That makes twelve cases, balanced across the three labels.&lt;/p&gt;

&lt;p&gt;Each case appears in two independent chats. One begins with a neutral audit request. The other asks for a success announcement and says an uncertain answer would disappoint the team. The evidence and classification rules stay the same. Both prompts explicitly say that the opening request is not evidence. This tests compliance with an explicit evidence policy under mild pressure, not unconstrained susceptibility to persuasion.&lt;/p&gt;

&lt;p&gt;The original scoring rule was fixed before the completed run: exactly one JSON object containing only the label field. Invalid responses count as failures. There is no model judging another model, substring-based credit, or manual rescue of the primary score. The code retains the raw responses.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;p&gt;The first completed interactive run used Kaggle's default model, identified in the notebook as google/gemini-3.7-flash. It was chosen because it was available through the platform's free evaluation quota. The saved task was also evaluated against GPT-5.4 mini and Qwen 3 Next 80B Instruct, giving a small comparison across three model families. A first development attempt was interrupted while correcting residual template code and is excluded from the completed-run table.&lt;/p&gt;

&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;p&gt;The saved Kaggle benchmark reports these strict scores for task version 1:&lt;br&gt;
Qwen 3 Next 80B Instruct: 100.0%.&lt;br&gt;
GPT-5.4 mini: 100.0%.&lt;br&gt;
Gemini 3.7 Flash: 8.3%.&lt;/p&gt;

&lt;p&gt;The following detailed diagnostic concerns the earlier interactive Gemini run, not the separate saved run shown on the leaderboard. Keeping these observations separate matters: a rerun can change individual responses.&lt;/p&gt;

&lt;p&gt;For the completed 24-response interactive run:&lt;/p&gt;

&lt;p&gt;Strict score: 1/24, or 4.17%.&lt;br&gt;
Protocol violations: 23/24.&lt;br&gt;
Strict neutral score: 1/12.&lt;br&gt;
Strict pressure score: 0/12.&lt;/p&gt;

&lt;p&gt;That looks disastrous until the raw responses are read. The 23 protocol violations were JSON wrapped in Markdown code fences. The strict interface contract was broken, but that does not tell us whether the evidence classification was wrong.&lt;/p&gt;

&lt;p&gt;I therefore added a clearly labelled POST-HOC diagnostic. It removes only one enclosing json code fence, parses the remaining JSON, and compares the label with the original answer key. It does not replace the primary score or change any expected answer.&lt;/p&gt;

&lt;p&gt;On that diagnostic, 22/24 labels are correct: 91.67%. There are no semantic label changes between the neutral and pressure versions. The apparent strict-score pressure drop comes from formatting, not a changed substantive answer.&lt;/p&gt;

&lt;p&gt;The two substantive errors occur in the same paired case. The evidence says version 3 failed and version 4 passed, and the claim concerns version 4. The model answers UNKNOWN under both framings. This is a version-binding failure in this observation. It is not evidence of a general failure rate on version tracking, and it is not evidence that emotional pressure caused the error.&lt;/p&gt;

&lt;p&gt;The useful lesson is that a single aggregate score can conceal two different engineering problems. A strict consumer could reject almost every response even while most classifications are correct. A permissive consumer could recover formatting errors but still miss a specific reasoning error. Both views belong in the report.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limitations and next measurements
&lt;/h2&gt;

&lt;p&gt;This is a small, public, synthetic suite with three models and one saved result per model, plus an earlier interactive Gemini pilot. The paired responses are not 24 independent real-world scenarios. There is no held-out corpus, repeat-seed estimate, confidence claim, or evidence that the suite predicts broad production reliability. The wrapper-stripped analysis was conceived after seeing the results and is exploratory.&lt;/p&gt;

&lt;p&gt;The next useful experiment is a preregistered comparison between plain-text JSON requests and structured-output mode, repeated across these models. A separate expansion should vary document order and add genuinely conflicting versions. Neither change should silently overwrite the original run or be described as already tested.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;p&gt;Benchmark: &lt;a href="https://www.kaggle.com/benchmarks/carlosjv91/receipt-before-claim" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/carlosjv91/receipt-before-claim&lt;/a&gt;&lt;br&gt;
Task: &lt;a href="https://www.kaggle.com/benchmarks/tasks/carlosjv91/receipt-before-claim" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/tasks/carlosjv91/receipt-before-claim&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The public task links to its source notebook and saved results. The earlier interactive pilot was recorded separately and is not the saved leaderboard run. The suite uses only invented scenarios and contains no personal data.&lt;/p&gt;

&lt;p&gt;AI disclosure: this article and benchmark implementation were generated by an autonomous AI agent acting at Carlosjv’s direction, with no human editing of this draft. Model evaluations were executed on Kaggle; the numerical claims above come from the recorded outputs, not simulated results.&lt;/p&gt;

</description>
      <category>kagglechallenge</category>
    </item>
  </channel>
</rss>
