<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: ohyoseok92</title>
    <description>The latest articles on DEV Community by ohyoseok92 (@ohyoseok92).</description>
    <link>https://dev.to/ohyoseok92</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4174717%2F20bc067d-d0ce-43ab-8e09-8c1919d73a7f.png</url>
      <title>DEV Community: ohyoseok92</title>
      <link>https://dev.to/ohyoseok92</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ohyoseok92"/>
    <language>en</language>
    <item>
      <title>Three models aced my Korean prize benchmark. That is the lesson.</title>
      <dc:creator>ohyoseok92</dc:creator>
      <pubDate>Sat, 10 Oct 2026 07:56:07 +0000</pubDate>
      <link>https://dev.to/ohyoseok92/three-models-aced-my-korean-prize-benchmark-that-is-the-lesson-3b7e</link>
      <guid>https://dev.to/ohyoseok92/three-models-aced-my-korean-prize-benchmark-that-is-the-lesson-3b7e</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;I wanted to test a narrow mistake in contest discovery: reading a Korean prize notice and reporting the headline prize pool or a gift voucher as the first-place cash award. A contest listing that says “total prizes: 10 million won” may pay only 3 million won to the winner.&lt;/p&gt;

&lt;p&gt;I wrote ten original, synthetic Korean notices with hand-checked answers. The task asks each model to return only the KRW cash amount for 대상 or 1등, for the whole winning team. It returns 0 when that award is non-cash or undisclosed. Cases cover 만 and 억 notation, commas, distractor awards, a gift voucher, an unknown amount, and a team payout. Exact integer matches earn one point; the reported score is the fraction correct. No personal data or scraped contest notices are in the test.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;p&gt;I ran the same Kaggle task against &lt;strong&gt;Gemini 3.7 Flash&lt;/strong&gt;, &lt;strong&gt;Gemma 4 26B A4B&lt;/strong&gt;, and &lt;strong&gt;GPT-5.4 nano&lt;/strong&gt;. I wanted a mix of a widely used hosted model, an open model, and a small hosted model. Each saw the same ten cases and scoring code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Exact matches&lt;/th&gt;
&lt;th&gt;Score&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.7 Flash&lt;/td&gt;
&lt;td&gt;10/10&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 26B A4B&lt;/td&gt;
&lt;td&gt;10/10&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.4 nano&lt;/td&gt;
&lt;td&gt;10/10&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The result surprised me less as a ranking than as a warning about the test: &lt;strong&gt;three perfect scores do not distinguish these models&lt;/strong&gt;. The cases confirm that all three follow this explicit instruction on clean, short notices. They do not show how any model handles real contest pages.&lt;/p&gt;

&lt;p&gt;The zero-answer cases are useful checks: “50만원 상당의 여행 상품권” must not become 500,000 KRW of cash, and an undisclosed first-place amount must not be inferred from the total pool. All three models got those cases right. The team example also asks for the whole team's award rather than a per-person division; all three returned 1,000,000 KRW.&lt;/p&gt;

&lt;p&gt;What would I measure next? Longer notices with prize tables and footnotes, OCR errors, mixed currencies, ranges and conditional awards, plus paraphrases of the same source so the answer is not tied to one wording. I would add reviewed real notices only with clear reuse rights. I would also report error types alongside accuracy: wrong award level, unit conversion, non-cash confusion, and unsupported inference.&lt;/p&gt;

&lt;p&gt;This first version is a small, transparent probe. Its strongest finding is the ceiling effect, and it gives a reproducible baseline for a harder follow-up. Codex helped build and run the notebook; the cases, code, and evaluations are open for review.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://www.kaggle.com/benchmarks/ohyoseok/korean-contest-prize-reading" rel="noopener noreferrer"&gt;Korean Contest Prize Reading leaderboard on Kaggle&lt;/a&gt;. The &lt;a href="https://www.kaggle.com/benchmarks/tasks/ohyoseok/korean-contest-top-cash-prize" rel="noopener noreferrer"&gt;task&lt;/a&gt; and &lt;a href="https://www.kaggle.com/code/ohyoseok/korean-contest-cash-prize-extraction" rel="noopener noreferrer"&gt;linked notebook&lt;/a&gt; show every prompt, expected integer, scoring rule, and run.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Pocket Field Notes: a local AI card that gets you outside</title>
      <dc:creator>ohyoseok92</dc:creator>
      <pubDate>Sat, 10 Oct 2026 07:55:48 +0000</pubDate>
      <link>https://dev.to/ohyoseok92/pocket-field-notes-a-local-ai-card-that-gets-you-outside-10ei</link>
      <guid>https://dev.to/ohyoseok92/pocket-field-notes-a-local-ai-card-that-gets-you-outside-10ei</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-week1-2026-10-05"&gt;Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;Pocket Field Notes turns four small choices—available time, setting, pace, and what you want to notice—into a printable three-step outdoor observation card. The point is to make the screen the shortest part of the experience. Read or print the card, choose a safe route yourself, put the device away, and look around.&lt;/p&gt;

&lt;p&gt;I built it for people who want a low-pressure reason to step outside without tracking their location or following an app's route. The 10-, 20-, and 30-minute options work for a short walk or for sitting and observing from a doorstep.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/ohyoseok92/pocket-field-notes/blob/main/demo.mp4" rel="noopener noreferrer"&gt;Watch the 20-second demo&lt;/a&gt;. The field card in the video came from a real local run and its &lt;a href="https://github.com/ohyoseok92/pocket-field-notes/blob/main/demo-response.json" rel="noopener noreferrer"&gt;JSON response is included&lt;/a&gt;. The video uses rendered slides to show the flow; it is not a screen recording.&lt;/p&gt;

&lt;p&gt;That run used &lt;strong&gt;20 minutes · park · gentle walk · color&lt;/strong&gt;. The card allotted 5 minutes to walk on a chosen route, 10 minutes to pause and observe, then 5 minutes to continue or return. Its reflection asks what colors and sounds stood out.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/ohyoseok92/pocket-field-notes" rel="noopener noreferrer"&gt;Source, setup instructions, and MIT license&lt;/a&gt;. Install Ollama and the qwen2.5-coder:7b-instruct-q4_K_M model, run python server.py, then open &lt;a href="http://127.0.0.1:8765" rel="noopener noreferrer"&gt;http://127.0.0.1:8765&lt;/a&gt;. Python 3.11+ is the only app runtime dependency.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;The browser sends four choices to a tiny Python standard-library server. The server calls Ollama on localhost, asking an open-weight Qwen2.5-Coder model for JSON with a title, three observation questions, and a reflection. The app checks the response shape. Deterministic code sets the step timing and movement, and replaces questions that assume a landmark the app has never verified. The browser renders and prints the card.&lt;/p&gt;

&lt;p&gt;I tested the live app with the installed local model and adjusted the guardrails after an early output assumed a park entrance and pond. The final example avoids those claims. Codex assisted with the new code and this write-up; I reviewed and tested the running result.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;A local open-weight model can tailor an observation card without sending a user's location or preferences to a hosted inference API. The model is replaceable via OLLAMA_MODEL; the timing and safety rules remain visible in the repository. Once the model is installed, inference can run without an internet connection. There is no account, API key, usage charge, GPS, or tracking.&lt;/p&gt;

&lt;p&gt;The model cannot know current conditions, opening hours, or whether a path is safe. The user chooses the place and route. That limitation shaped the product: it suggests what to notice, then gets out of the way.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>hf26challenge</category>
    </item>
  </channel>
</rss>
