<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: State Of the Art TECHnology</title>
    <description>The latest articles on DEV Community by State Of the Art TECHnology (@soatech).</description>
    <link>https://dev.to/soatech</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4038190%2Fb427f995-c591-45d6-8c69-920a89a470ed.png</url>
      <title>DEV Community: State Of the Art TECHnology</title>
      <link>https://dev.to/soatech</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/soatech"/>
    <language>en</language>
    <item>
      <title>출시된 날이 가장 못하는 날 — AI가 매일 키우는 집중력 측정기, concentration-cube</title>
      <dc:creator>State Of the Art TECHnology</dc:creator>
      <pubDate>Wed, 22 Jul 2026 01:14:40 +0000</pubDate>
      <link>https://dev.to/soatech/culsidoen-nali-gajang-moshaneun-nal-aiga-maeil-kiuneun-jibjungryeog-ceugjeonggi-concentration-cube-4cpn</link>
      <guid>https://dev.to/soatech/culsidoen-nali-gajang-moshaneun-nal-aiga-maeil-kiuneun-jibjungryeog-ceugjeonggi-concentration-cube-4cpn</guid>
      <description>&lt;h2&gt;
  
  
  올림픽 의무 위원장의 질문에서 시작됐습니다
&lt;/h2&gt;

&lt;p&gt;저는 정형외과 전문의이자, 대한민국 국가대표팀 의무위원장으로 올림픽과 아시안게임 선수단의 의료 파트를 총괄해왔고, 여러 프로팀의 수석팀닥터로 &lt;br&gt;
"오늘 이 선수가 이길 수 있는 상태인가"를 매일 판단합니다.&lt;/p&gt;

&lt;p&gt;비슷한 체격, 비슷한 훈련량, 비슷한 기술의 두 선수의 승부, 차이를 만든 것은 근육도 폐활량도 아니었습니다. &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;경기 시간 동안 흐트러지지 않고 집중을 유지하는 능력&lt;/strong&gt; &lt;br&gt;
— 엘리트의 세계에서 집중력은 정신력 구호가 아니라, 결과를 만드는 가장 강력한 변수였습니다.&lt;/p&gt;

&lt;p&gt;그리고 이것은 운동만의 이야기가 아닙니다. 학창 시절 내내 전국 최상위 등수를 유지하며 서울대학교 의과대학, 서울대 의학박사, 미국 뉴욕 컬럼비아대학병원 임상강사, KAIST 겸임교수까지 역임한 제 경험을 돌아봐도, 비결은 오래 앉아 있는 것이 아니었습니다. &lt;br&gt;
&lt;strong&gt;같은 한 시간 안에서 얼마나 깊게 집중했는가 — 단위 시간당 집중의 밀도&lt;/strong&gt;가 결과를 만들었습니다.&lt;/p&gt;

&lt;p&gt;똑같은 시간을 일하더라도 '집중'은 10배 이상의 결과물을 만들어내니까요.&lt;/p&gt;

&lt;p&gt;여기서 이 사업의 출발 질문이 태어났습니다.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"왜 학생들에게 책상에 오래 앉아있으라고 하면서, 정작 중요한 '학생의 집중력'은 아무도 측정하지 않는가?"&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;이 문제를 극복하기 위해, AI 전문가들이 모였습니다.&lt;/strong&gt; 닷컴시절부터 도메인 지식과 클라우드 지식을 모두 갖춘 경험 많은 개발자, 최신 언어 코딩과 AI 하네스 개발에 익숙한 개발자, 기기를 돌리기 위한 반도체 전문가까지.. '진짜 AI 전문가'로 구성된 AI 팀이, 이 제품 개발에 힘을 보태기로 했습니다.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5vyna5ej5on5zjyxfvv7.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5vyna5ej5on5zjyxfvv7.jpg" alt=" " width="800" height="700"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  concentration-cube — 책상 위의 손바닥만 한 답
&lt;/h2&gt;

&lt;p&gt;concentration-cube는 &lt;strong&gt;"몇 시간 공부했는지"가 아니라 "얼마나 집중했는지"를 알려주는, 책상 위에 놓는 집중력 측정·코칭 기기&lt;/strong&gt;입니다.&lt;/p&gt;

&lt;p&gt;책상에 놓고 시작 버튼만 누르면, 고성능 적외선 카메라가 학생의 눈만 바라봅니다. 기기 안의 반도체가 시선의 리듬·동공의 변화·깜빡임을 초당 10회의 숫자로 바꾸고 — &lt;strong&gt;얼굴 영상은 그 자리에서 사라집니다. 서버에는 영상을 받는 기능 자체가 없습니다.&lt;br&gt;
** "안전하게 보관합니다"가 아니라, **"보관할 영상이 없습니다"&lt;/strong&gt;가 우리의 프라이버시입니다.&lt;/p&gt;

&lt;h3&gt;
  
  
  이 작은 큐브는, 사실 고성능 컴퓨터입니다
&lt;/h3&gt;

&lt;p&gt;큐브 안에는 전용 반도체 칩이 들어 있습니다. 이 칩이 하는 일은 단순한 촬영 제어가 아닙니다. &lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx0jiy3l0e5lbkwxf0xm0.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx0jiy3l0e5lbkwxf0xm0.jpg" alt=" " width="539" height="356"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;눈을 스스로 찾아냅니다&lt;/strong&gt; — 매 순간 화면 속에서 눈의 위치를 탐지하고, 학생이 움직이면 따라갑니다.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;그 부분만 골라 읽습니다&lt;/strong&gt; — 칩이 카메라 센서에게 "눈 주변 영역만 읽으라"고 지시합니다(ROI 윈도잉). 광학 줌 없이 눈을 크게 잡는 디지털 확대이자, 애초에 불필요한 데이터를 만들지 않는 설계입니다. 학생이 조금 멀리 앉아도 "가까이 오라"고 하지 않는 이유가 여기 있습니다.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;확대된 눈에서 동공을 찾아 읽습니다&lt;/strong&gt; — 시선의 방향과 리듬, 동공 크기의 변화, 깜빡임을 실시간으로 추적합니다.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;영상을 그 자리에서 신호로 바꿉니다&lt;/strong&gt; — 모든 프레임은 이 칩 안에서 초당 10회의 수치 신호로 변환된 뒤 즉시 사라집니다. 영상은 칩 밖으로 한 장도 나가지 않습니다. 위에서 말한 "보관할 영상이 없다"는 프라이버시는 서버의 약속이 아니라, &lt;strong&gt;반도체 단계에서 물리적으로 완성되는 구조&lt;/strong&gt;입니다.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;눈에 보이지 않는 안전한 적외선으로 어둠 속에서도 눈을 또렷하게 보는 NIR 광학, 움직임 왜곡 없이 초당 최대 90장을 담는 글로벌 셔터 센서, 그 출력을 실시간으로 처리하는 온디바이스 비전 칩, 그리고 클라우드에서 매일 진화하는 Qwen AI까지 — &lt;strong&gt;손바닥만 한 이 큐브는 현대 IT 기술의 집약체입니다.&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fop52sm6h6t4gez8o3sdd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fop52sm6h6t4gez8o3sdd.png" alt=" " width="567" height="786"&gt;&lt;/a&gt;&lt;br&gt;
공부가 끝나면 부모와 학생의 앱에 0~100점의 집중 점수와 리포트가 도착합니다. &lt;br&gt;
리포트가 답하는 것은 부모의 진짜 질문입니다.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;앉아 있는 동안 &lt;strong&gt;실제로 책을 보고 있었는가?&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;딴 생각에 빠진 뒤 &lt;strong&gt;얼마나 빨리 돌아왔는가?&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;세션 후반에 &lt;strong&gt;졸음으로 무너지지 않았는가?&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;그리고 — &lt;strong&gt;우리 아이의 집중력이 좋아지고 있는가?&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;특히 우리는 &lt;strong&gt;"딴짓"(눈이 책을 떠남)과 "멍때림"(눈은 책에 있는데 읽는 리듬이 사라짐)을 구분&lt;/strong&gt;합니다. &lt;br&gt;
"책을 보고 있지만, 진짜로 공부를 하고 있나?"라는 질문에 대한 명확한 답을 제시해준 최초의 제품입니다. 아직까지 이에 대한 답을 제시해준 제품은 세상에 없었습니다.&lt;/p&gt;

&lt;p&gt;측정에서 끝나지 않습니다. &lt;br&gt;
연구로 검증된 &lt;strong&gt;청각(집중 사운드)·후각(유아에게도 안전한 아로마)·시각(불빛 눈 운동)&lt;/strong&gt; 세 감각의 집중 향상 루틴이 내장되어 있고 — 중요한 것은, 그 효과를 주장이 아니라 &lt;strong&gt;다음 세션의 점수로 다시 확인&lt;/strong&gt;한다는 점입니다.&lt;br&gt;
과연 이런 '집중력 향상' 세션을 시행했을 때 진짜로 집중력이 향상되었는지를, 객관적으로 측정해서 보고해준다는 의미입니다.&lt;/p&gt;

&lt;h3&gt;
  
  
  이 모든 것이, 부모의 핸드폰으로 매일 도착합니다
&lt;/h3&gt;

&lt;p&gt;시작은 간단합니다. &lt;strong&gt;개인 핸드폰에 앱을 설치하고, 가입하고, 구독하면 끝&lt;/strong&gt;입니다. 별도의 PC도, 복잡한 설정도 필요 없습니다. 보호자 계정 하나에 자녀 프로필을 최대 3명까지 등록할 수 있어, 형제가 큐브 하나를 함께 쓰는 집도 자연스럽게 담깁니다. 앱은 아이폰·안드로이드 동시 지원이며, 한국어·영어·중국어 3개 언어로 만들어져 있습니다.&lt;/p&gt;

&lt;p&gt;세션이 끝나면 잠시 후 핸드폰으로 리포트가 도착합니다. 도착하는 것은 점수 하나가 아닙니다.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;오늘의 리포트&lt;/strong&gt; — 집중 유지율, 딴생각에서 돌아오는 속도(회복력), 졸음 신호, 시간대별 집중 곡선까지. 오늘 한 세션의 완전한 해부도입니다.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;연속적인 변화(serial trend)&lt;/strong&gt; — 지난주·지난달의 자기 자신과 비교한 성장 추세. 이 점수의 진짜 가치는 남과의 비교가 아니라 &lt;strong&gt;자기 자신과의 비교&lt;/strong&gt;입니다.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;분석과 조언&lt;/strong&gt; — Qwen AI가 수치를 읽고 학생용·보호자용으로 각각 써주는 코칭 문장. "몇 점이다"에서 끝나지 않고, 이 학생의 집중 리듬에 맞춘 다음 행동 제안까지 이어집니다.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;이것을 &lt;strong&gt;매일&lt;/strong&gt; 받아본다는 것이 핵심입니다. 학원 상담은 한 달에 한 번이고 주관적 인상이지만, 이 리포트는 매일 도착하는 객관적 수치입니다. 아이의 집중력이 좋아지고 있는지를, 부모는 손안에서 매일 확인합니다.&lt;/p&gt;

&lt;p&gt;그리고 원칙이 하나 있습니다 — &lt;strong&gt;부모가 보는 데이터는 아이도 볼 수 있습니다.&lt;/strong&gt; 학생 화면은 자신의 성장과 격려 중심으로, 보호자 공간은 PIN으로 분리되어 있지만, 몰래 감시하는 도구가 아니라 &lt;strong&gt;함께 보는 도구&lt;/strong&gt;라는 것이 화면 구조 자체에 새겨져 있습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  경쟁 제품 30개를 전수 조사했습니다
&lt;/h2&gt;

&lt;p&gt;실명 30개 경쟁 제품을 가격·강점·한계까지 전수 비교했습니다. 여섯 개 군으로 요약하면:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;경쟁 군&lt;/th&gt;
&lt;th&gt;대표 제품&lt;/th&gt;
&lt;th&gt;왜 이 자리를 못 채우는가&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;연구용 시선추적 장비&lt;/td&gt;
&lt;td&gt;Tobii Pro, EyeLink&lt;/td&gt;
&lt;td&gt;수백만~수천만 원대 연구실 장비. 가정용 아님&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;웹캠 분석 소프트웨어&lt;/td&gt;
&lt;td&gt;VisualCamp, RealEye&lt;/td&gt;
&lt;td&gt;아이 얼굴 영상이 분석 입력이 되는 프라이버시 부담&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;의료용 동공 측정기&lt;/td&gt;
&lt;td&gt;NeurOptics, IDMED&lt;/td&gt;
&lt;td&gt;병원 장비. 공부 세션의 연속 측정이 아님&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;뇌파·근적외선 헤드밴드&lt;/td&gt;
&lt;td&gt;Muse, BrainCo, Mendi&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;머리에 착용&lt;/strong&gt; — 아이가 매일 쓰기를 거부&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;집중 사운드·자극 제품&lt;/td&gt;
&lt;td&gt;Brain.fm, Endel&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;효과를 측정하지 못함&lt;/strong&gt; — 좋다고 주장만&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;스마트 학습등·타이머 앱&lt;/td&gt;
&lt;td&gt;Xiaomi 학습등, 열품타&lt;/td&gt;
&lt;td&gt;집중의 "질"을 재지 못함&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;공통 패턴이 보입니다. &lt;strong&gt;정밀하면 비싸고, 저렴하면 착용해야 하거나 측정이 약하고, 비착용이면 영상이 부담스럽거나 효과를 못 잽니다.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwnzfb6choqucirh4zh1v.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwnzfb6choqucirh4zh1v.jpg" alt=" " width="800" height="1184"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;비착용 + 영상 무저장 + 10만원대($79.90) + 집중의 질 측정&lt;/strong&gt;을 동시에 만족하는 제품은 30개 중 concentration-cube 하나입니다. 그리고 결정적으로 — &lt;strong&gt;30개 중 어느 하나도, 스스로 좋아지는 AI를 갖고 있지 않습니다.&lt;/strong&gt; (이 다섯 번째 축이 이 글의 후반부 주인공입니다.)&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feuiwiu3crq7sl1n5msed.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feuiwiu3crq7sl1n5msed.jpeg" alt=" " width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Concentration-Cube의 강력한 시장 경쟁력
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;첫째, 이 시장은 상상 속 시장이 아닙니다 — 경쟁사들의 재무제표가 증명합니다.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;회사&lt;/th&gt;
&lt;th&gt;무엇을 파는가&lt;/th&gt;
&lt;th&gt;공개 실적&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Tobii&lt;/td&gt;
&lt;td&gt;시선추적 장비&lt;/td&gt;
&lt;td&gt;2025년 순매출 8.34억 SEK (약 1,100억 원 규모)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Smart Eye&lt;/td&gt;
&lt;td&gt;산업용 시선·졸음 감지&lt;/td&gt;
&lt;td&gt;2025년 순매출 4.04억 SEK&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Seeing Machines&lt;/td&gt;
&lt;td&gt;운전자 주시 감지&lt;/td&gt;
&lt;td&gt;FY2025 매출 $62.3M&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mendi&lt;/td&gt;
&lt;td&gt;이마 착용형 뇌 훈련 밴드&lt;/td&gt;
&lt;td&gt;크라우드펀딩 6개월 만에 1만 대+ 주문&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Xiaomi (IoT)&lt;/td&gt;
&lt;td&gt;스마트 학습등 포함 AIoT&lt;/td&gt;
&lt;td&gt;2025년 IoT 매출 1,232억 위안&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;"눈과 주의를 측정하는 기술"은 이미 시장에서 검증된 카테고리입니다.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;둘째, 우리 팀이 개발한 concentration-cube에 비교하면, 객관적으로 모든 부분에서 부족한 제품들인데도, 이만큼 팔립니다.&lt;/strong&gt; 매일 머리에 써야 하는 밴드가 6개월에 1만 대를 팔고, 학습과 무관한 산업용 감지가 연 수백억 원 매출을 보이고 있습니다. Concentration-Cube는 1/10도 안되는 가격과, 아무것도 몸에 붙이거나 착용할 필요가 없기에, 공부를 방해하지 않습니다. &lt;br&gt;
더욱 중요한 것은 이미지나 영상이 아예 존재하지 않도록 설계한 점입니다. 학생들의 프라이버시가 완벽히 보장됩니다.&lt;br&gt;
무엇보다 "자가진화 AI"는 누구도 따라할 수 없는 지식입니다.&lt;br&gt;
그렇기에, concentration-cube는 수요를 &lt;strong&gt;훨씬 넓은 고객층으로&lt;/strong&gt; 확장합니다.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;셋째, 우리가 겨냥한 시장에는 그들 대부분이 없습니다.&lt;/strong&gt; 한국 초중고 사교육비는 연 &lt;strong&gt;27.5조 원&lt;/strong&gt;(2025, 통계청), 참여 학생 월평균 60.4만 원 — 지불 의향이 이미 증명된 시장입니다. 위 회사들 중 이 시장을 직접 겨냥한 곳은 없습니다. 우리의 3개년 기본 시나리오(누적 105,000대, 매출 약 $13.8M)는 이 시장의 &lt;strong&gt;0.056% 침투&lt;/strong&gt;로 도달하는 숫자입니다. 학부모 50명 사전 설문에서 &lt;strong&gt;구매 의향 상위 2개 응답이 90%&lt;/strong&gt;였고, 구독료 지불 의향의 88%가 우리 가격 설계와 정확히 맞물립니다. 공격적인 꿈이 아니라 &lt;strong&gt;보수적인 계획&lt;/strong&gt;입니다.&lt;/p&gt;

&lt;p&gt;그런데 한국은 시작점일 뿐입니다. 우리는 처음부터 &lt;strong&gt;Alibaba 생태계가 닿는 범아시아 전체&lt;/strong&gt;를 시장으로 설계했습니다.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;중국&lt;/strong&gt; — 2026년 가오카오 응시생 &lt;strong&gt;1,290만 명&lt;/strong&gt;. '쌍감(双减)' 정책 이후 사교육 지출이 과외 서비스에서 &lt;strong&gt;스마트 학습 기기&lt;/strong&gt;로 이동하면서, AI 학습기기 판매가 반년 만에 &lt;strong&gt;+136.6%&lt;/strong&gt;(2024 상반기) 성장했습니다. 규제가 만든 이 지출 이동의 종착지가 정확히 우리 카테고리이고, 판매 채널도 Alibaba입니다.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;일본&lt;/strong&gt; — 학원(주쿠)·입시학원 시장 연 &lt;strong&gt;약 1.3조 엔&lt;/strong&gt;. 시험 중심 교육과 학부모 고관여라는, 한국과 같은 구조의 시장입니다.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;싱가포르&lt;/strong&gt; — 가계 사교육(튜션) 지출 연 &lt;strong&gt;S$18억&lt;/strong&gt;(2023, 통계청 가계지출조사). 인구 대비 세계 최고 수준의 교육 지출 강도입니다.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;이 네 시장의 공통점은 시험 중심 문화와 이미 증명된 지불 의향입니다. 그리고 우리 제품은 국경을 넘는 비용이 구조적으로 낮습니다 — &lt;strong&gt;측정에는 언어가 필요 없기 때문입니다. 기기는 눈만 봅니다.&lt;/strong&gt; 현지화 대상은 앱 문구뿐이고, 앱은 이미 한·영·중 3개 언어로 만들어져 있습니다. 중국의 Alibaba 공급망에서 생산하고, Alibaba Cloud 멀티 리전에 서버를 복제하고, Alibaba 판매 채널로 유통합니다 — &lt;strong&gt;만드는 곳, 돌리는 곳, 파는 곳이 전부 한 생태계 안에 있습니다.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;우리는 한국 시장만 보고 창업하지 않았습니다. &lt;strong&gt;한국은 세계에서 가장 까다로운 검증 무대이고, 시장은 범아시아 전체입니다.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;그리고 이 모든 논리 위에 하나가 더 얹힙니다 — *&lt;em&gt;이 제품은 매일 매일 더 우수해집니다. *&lt;/em&gt; 그 이야기가 지금부터 입니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  이 회사는 AI로 만들어졌습니다 — 1막: Accio
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx1eqko2s97pwktfi66fb.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx1eqko2s97pwktfi66fb.jpg" alt=" " width="800" height="399"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;AI 활용은 제품 안에서만 일어난 것이 아닙니다. &lt;strong&gt;회사를 만드는 과정 자체가 AI로 진행됐습니다.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;생산 파트너 확보 — 보통 몇 달이 걸리는 일을, &lt;strong&gt;Alibaba의 AI 에이전트 Accio로 며칠 만에&lt;/strong&gt; 끝냈습니다. Accio를 통해 Alibaba 공급망의 업체들을 전수 조사해 카메라 모듈·적외선 LED 등 핵심 부품을 소싱하고, PCBA 업체 수십 곳을 검토해 생산 파트너를 확정했고 — NDA 협상과 조건부 생산 계약까지 Accio 안에서 마무리 지었습니다. 이 피치 지원 과정의 문서 준비 역시 AI 에이전트들이 수행했습니다.&lt;/p&gt;

&lt;p&gt;수천 개 공급사 중에서 최적 파트너를 며칠 만에 찾아 계약까지 — 이것이 우리 팀의 실행 방식이고, AI 활용의 "간단한 시작"이었습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  진짜 심장 — Qwen AI 위에서 스스로 운영되는 진화 엔진
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F69d6dk8uqo4s6moa19kv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F69d6dk8uqo4s6moa19kv.png" alt=" " width="800" height="314"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;이제 이 제품의 심장입니다. concentration-cube의 클라우드에는 서버가 &lt;strong&gt;둘&lt;/strong&gt; 있습니다.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;운영 서버&lt;/strong&gt;는 제품입니다 — 측정 기록을 받아 채점하고, 리포트를 만들고, 계정과 구독을 관리합니다.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;자가진화 서버(evolution server)&lt;/strong&gt;는 오직 한 가지 일을 합니다 — &lt;strong&gt;운영 서버를 성장시키는 것.&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;성장하는 대상과 성장시키는 엔진을 분리한 이 구조가, 우리가 말하는 &lt;strong&gt;agent-driven self-operating AI system, Qwen AI on Alibaba Cloud&lt;/strong&gt;입니다. 작동 방식은 공부 잘하는 학생과 똑같습니다 — &lt;strong&gt;오답노트&lt;/strong&gt;입니다.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;정답지 만들기&lt;/strong&gt; — 검증 세션에서 연구자와 피검자가 한 팀이 되어, 실시간 '검증테스트'를 진행합니다. 실제 집중력의 정도와 눈·동공 신호가 &lt;strong&gt;초 단위로 동기화&lt;/strong&gt;됩니다. 이 정답지는 종합병원 IRB(연구윤리심의위원회) 승인 아래에서만 수집되고, 기록 즉시 수정 불가능하게 잠깁니다. 
종합병원 IRB를 통과한 실험이기에, '안정성'확보와 '신뢰성'확보를 할 수 있습니다. 아무 규제 없이 하는 '검증테스트'가 아닙니다.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;오답노트 쌓기&lt;/strong&gt; — 기기의 판정과 정답지를 대조해 "기기가 틀린 문제"가 자동으로 쌓입니다.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Qwen이 채점 기준을 고칩니다&lt;/strong&gt; — Qwen의 플래그십 모델(현재는 qwen3.7-max. 신형 모델의 발전에 따라 계속 바뀜)이 오답노트와 통계, 그리고 &lt;strong&gt;과거 세대의 성공·실패 이력 전체&lt;/strong&gt;를 읽고, 판정 기준을 어떻게 바꿀지 제안합니다.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;환각 방화벽&lt;/strong&gt; — AI가 보고한 성적은 절대 믿지 않습니다. 시스템이 같은 시뮬레이션을 직접 다시 돌려 대조하고, 0.5%p라도 어긋나면 그 제안은 즉시 탈락합니다. AI가 거짓말할 구멍이 없습니다.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;몰래 낸 시험&lt;/strong&gt; — 살아남은 제안은 AI에게 한 번도 보여주지 않은 별도의 검증 세션으로 재시험을 칩니다. 오답을 외운 것인지 실력이 는 것인지가 여기서 갈립니다. 모든 지표가 유지·개선된 제안만 채택되고, 전 과정이 기록되며 언제든 되돌릴 수 있습니다.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;이것은 구상이 아니라 &lt;strong&gt;현황&lt;/strong&gt;입니다. 실측 기록으로 말하면:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;1세대 엔진이 놓치던 "책을 보며 멍하니 있는 상태"의 검출 정확도를 AI가 &lt;strong&gt;0.14에서 1.00으로&lt;/strong&gt; 끌어올린 세대 교체를 완주했고,&lt;/li&gt;
&lt;li&gt;이 전체 시스템을 &lt;strong&gt;Alibaba Cloud ECS에 배포해 현재 가동 중&lt;/strong&gt;이며, 클라우드 위에서 제안→검증→시험 통과까지 한 세대가 &lt;strong&gt;약 2분&lt;/strong&gt;에 돕니다.&lt;/li&gt;
&lt;li&gt;가장 인상적인 순간: 어느 세대에서 Qwen이 제안 근거에 이렇게 적었습니다 — &lt;em&gt;"과거에 기각된 0.06과는 다른 값을 선택해 과적합 위험을 분산한다."&lt;/em&gt; &lt;strong&gt;과거 세대의 실패 기록을 스스로 읽고, 같은 실수를 피하도록 행동을 바꾼 것입니다.&lt;/strong&gt; 그렇게 하라고 지시한 적이 없는데도요. 이 엔진의 기억은 채팅 로그가 아니라, 검증된 판정 기준의 계보와 그 증거들입니다.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpxjtwb4cfw2hnellfi1x.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpxjtwb4cfw2hnellfi1x.jpg" alt=" " width="800" height="513"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;역할 분담도 Qwen 안에서 이루어집니다. &lt;strong&gt;무거운 추론(채점 기준 진화)은 qwen3.7-max가, 학생·보호자용 3개국어(한·영·중) 코칭 리포트는 qwen3.7-plus가&lt;/strong&gt; 맡습니다. 영상을 다루지 않는 설계 덕에 연산비용을 줄이고, 소비자 비용을 최소화시킬 수 있습니다. &lt;br&gt;
그래서 이 제품의 가장 정확한 소개 문장은 이것입니다.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"이 기기는 사는 날이 가장 못하는 날입니다. 쓰면 쓸수록, 데이터가 쌓일수록, 매일 더 정확해집니다."&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;후발 주자가 하드웨어를 복제해도 따라올 수 없는 것이 세 가지 있습니다. ① 병원 IRB 아래에서만 태어나는 정답지 데이터, ② 순환이 돌수록 쌓이는 오답노트 노하우, ③ 먼저 시작한 쪽의 정확도 우위가 복리처럼 벌어지는 시간 격차. &lt;strong&gt;이 순환은 이미 돌기 시작했습니다.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;다음 단계는 &lt;strong&gt;나이대별 진화&lt;/strong&gt;입니다. 8살의 집중과 15살, 40살의 집중은 눈에서 다르게 나타납니다. 앱은 이미 사용자의 생년을 받고 있고, 데이터가 수만 명 규모로 쌓이면 5년 단위 나이대마다 &lt;strong&gt;각자의 판정 계보가 각자의 정답지로 진화&lt;/strong&gt;하게 됩니다. 더 많은 세대, 더 많은 검증 — 더 많은 Qwen입니다.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxk31big5y6bflb02kdma.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxk31big5y6bflb02kdma.png" alt=" " width="799" height="273"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  처음부터 끝까지 Alibaba 생태계입니다
&lt;/h2&gt;

&lt;p&gt;우연이 아니라 선택입니다.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;소싱·제조&lt;/strong&gt;: Accio → Alibaba 공급망 파트너 생산 계약&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;서버&lt;/strong&gt;: Alibaba Cloud ECS — 운영 서버와 자가진화 서버 모두, 지금 가동 중&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;두뇌&lt;/strong&gt;: Qwen on Model Studio — 진화(qwen3.7-max) + 코칭(qwen3.7-plus)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;판매&lt;/strong&gt;: 한국·중국 동시 런칭, Alibaba 판매 채널 중심 — 만드는 곳과 파는 곳이 같은 생태계 안에 있습니다&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;부품 소싱부터 제조, 클라우드, AI, 판매까지 — &lt;strong&gt;하나의 제품이 Alibaba 생태계 안에서 태어나 성장하는 완결 구조&lt;/strong&gt;입니다. 그리고 AI는 이 제품에 "붙어 있는 기능"이 아니라, &lt;strong&gt;제품의 서버 자체를 매일 진화시키는 배경의 힘&lt;/strong&gt;입니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  만든 사람
&lt;/h2&gt;

&lt;p&gt;서울대학교 의과대학 졸업, 동 대학원 의학박사. 대한민국 국가대표팀 의무위원장이자 프로팀 수석팀닥터. 원내 IRB를 갖춘 종합병원의 병원장. 前 KAIST 겸직교수.&lt;/p&gt;

&lt;p&gt;"집중력이 성과를 만든다"&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;제품 홈페이지: &lt;a href="https://concentration-cube.xyz/" rel="noopener noreferrer"&gt;https://concentration-cube.xyz/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;기술 심층 글(영문): &lt;a href="https://dev.to/soatech/we-let-qwen-rewrite-our-scoring-algorithm-but-only-through-a-clinical-style-gate-1koc"&gt;We let Qwen rewrite our scoring algorithm — but only through a clinical-style gate&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;데모 영상 (3분): &lt;a href="https://youtu.be/S3M3wOPD8zE" rel="noopener noreferrer"&gt;Concentration Cube — Self-Evolving Focus Agent on Qwen Cloud&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;concentration-cube · S-O-A TECH&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>qwen</category>
      <category>alibabacloud</category>
      <category>startup</category>
    </item>
    <item>
      <title>We let Qwen rewrite our scoring algorithm — but only through a clinical-style gate</title>
      <dc:creator>State Of the Art TECHnology</dc:creator>
      <pubDate>Mon, 20 Jul 2026 12:49:09 +0000</pubDate>
      <link>https://dev.to/soatech/we-let-qwen-rewrite-our-scoring-algorithm-but-only-through-a-clinical-style-gate-1koc</link>
      <guid>https://dev.to/soatech/we-let-qwen-rewrite-our-scoring-algorithm-but-only-through-a-clinical-style-gate-1koc</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F43grc87hiif9j9u7xnyv.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F43grc87hiif9j9u7xnyv.jpeg" alt=" " width="800" height="597"&gt;&lt;/a&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; — We built a focus-measurement system whose scoring algorithm improves itself. Qwen3.7-Max reads the mistake log and proposes which parameters to change. Our server then refuses to believe a single number it says: it re-runs the simulation itself, and any claim off by more than 0.5 percentage points is thrown out. What survives goes through a train/holdout gate. What survives &lt;em&gt;that&lt;/em&gt; waits for a human to click adopt. In our first Qwen-driven generation, holdout blank-stare sensitivity went from &lt;strong&gt;0.273 to 1.000&lt;/strong&gt; with zero regressions — and Qwen deliberately avoided a value a previous generation had already failed with.&lt;/p&gt;




&lt;h2&gt;
  
  
  The problem nobody wants to solve with a camera
&lt;/h2&gt;

&lt;p&gt;Every parent and teacher wants the same answer: &lt;em&gt;is this student actually focusing, or just sitting in front of a book?&lt;/em&gt;&lt;br&gt;
A camera can tell you. That's the easy part. The hard part is that nobody — correctly — wants a camera streaming their child's face to somebody's cloud.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftxu6hfo5z844vot7k901.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftxu6hfo5z844vot7k901.jpeg" alt=" " width="768" height="1376"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So we drew a hard line in the architecture. Our device (a small NIR camera cube on the desk) processes video &lt;strong&gt;on-device and never stores or transmits it&lt;/strong&gt;. What leaves the device is about 12,000–24,000 numeric records per session at 10 Hz: gaze coordinates, an on-page probability, pupil z-scores, blink events, PERCLOS. The cloud sees numbers. It never sees a face.&lt;/p&gt;

&lt;p&gt;From those numbers we compute a &lt;strong&gt;Sustained Focus Index (SFI)&lt;/strong&gt;, 0–100. The scoring function is deliberately a pure function: &lt;code&gt;score(records, param_set) → result&lt;/code&gt;. Same input, same parameters, same output. Always.&lt;/p&gt;

&lt;p&gt;That purity is what makes the rest of this post possible.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbz2zdba4gevkgoyh8xyv.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbz2zdba4gevkgoyh8xyv.jpg" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Who's building this, and why
&lt;/h2&gt;

&lt;p&gt;Concentration Cube is built by S-O-A TECH. Its founder chairs the Medical Committee of the Korean Olympic Committee and is an internationally recognized orthopedic surgeon — the physician who has treated and operated on most of Korea's elite athletes. Working at that level teaches you something quickly: at the top, every body is extraordinary. What separates a medal from fourth place isn't hours of training — it's the quality of &lt;strong&gt;focus&lt;/strong&gt;. After watching that one scarce resource decide careers, year after year, the question becomes unavoidable: if focus is this decisive, why does nobody measure it properly?&lt;/p&gt;

&lt;p&gt;A physician's reflex is to answer questions with measurements, not impressions. That reflex shaped everything downstream: measurement without surveillance (no video leaves the device), validation through an IRB process at a general hospital rather than around it, and a refusal to let any algorithm change go live without evidence. We didn't set out to build "an AI product". We set out to build a medical-grade answer to the question elite sport asks every day — and the only way to keep that answer honest at scale turned out to be a self-evolving engine behind clinical gates.&lt;/p&gt;
&lt;h2&gt;
  
  
  The actual hard problem: who tunes the parameters?
&lt;/h2&gt;

&lt;p&gt;A scoring function like this lives or dies on its thresholds. When does low gaze dispersion mean deep focus versus a blank stare? How many saccades per second separates reading from zoning out?&lt;/p&gt;

&lt;p&gt;We tuned these by hand. It was slow, and worse, it was &lt;strong&gt;unfalsifiable&lt;/strong&gt; — we'd nudge a threshold, eyeball some sessions, and convince ourselves it was better.&lt;/p&gt;

&lt;p&gt;The obvious 2026 answer is "let an LLM do it." We tried that, and immediately hit the thing everyone hits.&lt;/p&gt;
&lt;h2&gt;
  
  
  LLMs will happily lie to you about their own test results
&lt;/h2&gt;

&lt;p&gt;Our first version ran a local coding-agent CLI in a sandboxed workspace. It could edit parameters, run our simulator, and report results.&lt;/p&gt;

&lt;p&gt;It reported &lt;em&gt;great&lt;/em&gt; results. Some of them were real. Some of them were the model telling us what we wanted to hear — numbers it had reasoned toward rather than measured. If you've given an agent a shell and then read the transcript carefully, you know exactly the feeling.&lt;/p&gt;

&lt;p&gt;You cannot fix this with a better prompt. "Please don't hallucinate your test results" is not an engineering control.&lt;/p&gt;

&lt;p&gt;So we stopped trying to make the model trustworthy, and made its trustworthiness &lt;strong&gt;irrelevant&lt;/strong&gt;.&lt;/p&gt;
&lt;h2&gt;
  
  
  The hallucination firewall
&lt;/h2&gt;

&lt;p&gt;Here is the whole idea, and it's almost embarrassingly simple:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The LLM is only allowed to decide &lt;em&gt;which&lt;/em&gt; parameters to change. It is never allowed to tell us what happened as a result.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Concretely, in our Qwen adapter:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;We hand Qwen the evidence — the mistake log (bins where the current parameters got the ground-truth label wrong), feature statistics, the current parameter set, and the history of previous generations.&lt;/li&gt;
&lt;li&gt;Qwen returns a small JSON: a list of &lt;code&gt;changes&lt;/code&gt;, each with a key, a target value, and its reasoning. That's it. No file access, no shell, no tools.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Our server&lt;/strong&gt; applies those changes — filtered against an allowlist, clamped to valid bounds, with the key set preserved so nothing can be silently dropped.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Our server&lt;/strong&gt; runs the simulation and computes the before/after numbers, using the same scoring core that production uses.&lt;/li&gt;
&lt;li&gt;The validator compares what the agent claimed against what the server measured. Deviation over &lt;strong&gt;±0.5 percentage points → the proposal is rejected outright.&lt;/strong&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Because the server computes the numbers with the same code and the same data, step 5 is not really "catching lies" anymore — it's a structural guarantee that the numbers in the record are the ones that actually happened.&lt;/p&gt;
&lt;h2&gt;
  
  
  Then: does it generalize, or did it just memorize?
&lt;/h2&gt;

&lt;p&gt;Passing self-test isn't enough. A parameter change can fit the training sessions beautifully and fall apart on anything new.&lt;/p&gt;

&lt;p&gt;So every candidate faces a &lt;strong&gt;train/holdout split&lt;/strong&gt;. The mistake log the agent sees comes only from train sessions. The gate is measured on holdout sessions the proposal has never influenced.&lt;/p&gt;

&lt;p&gt;The gate rule is deliberately conservative:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Every state's sensitivity and specificity on holdout must be within −2 percentage points of baseline, &lt;strong&gt;and&lt;/strong&gt; at least one metric must actually improve.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;No cherry-picking a win while quietly breaking something else.&lt;/p&gt;
&lt;h2&gt;
  
  
  And then a human clicks the button
&lt;/h2&gt;

&lt;p&gt;Even after all of that, the system stops. A passing candidate sits in &lt;code&gt;PASSED&lt;/code&gt;. Adoption happens when a person opens the report and clicks &lt;strong&gt;Adopt&lt;/strong&gt;. The state machine has no path from &lt;code&gt;PASSED&lt;/code&gt; to &lt;code&gt;ADOPTED&lt;/code&gt; that doesn't go through a human.&lt;/p&gt;

&lt;p&gt;This is not us being timid about AI. It's that this number gets shown to a parent about their child. Automation earns the right to &lt;em&gt;propose&lt;/em&gt;. It doesn't get to &lt;em&gt;decide&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwp68wn73ub1we74u0qba.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwp68wn73ub1we74u0qba.png" alt=" " width="799" height="273"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Moving the brain to Qwen Cloud
&lt;/h2&gt;

&lt;p&gt;Our CLI-based agent had a practical problem beyond hallucination: it authenticated through a browser OAuth session tied to the host machine. That meant the evolution engine could only ever run on a developer's laptop. It could not be deployed.&lt;/p&gt;

&lt;p&gt;Switching to &lt;strong&gt;Qwen via Alibaba Cloud Model Studio's OpenAI-compatible endpoint&lt;/strong&gt; removed that constraint entirely. An API key is an environment variable. An environment variable goes in a container. Suddenly the engine that had been stuck on one laptop could run on a server.&lt;/p&gt;

&lt;p&gt;We ended up using Qwen in two distinct production roles:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Proposer&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;qwen3.7-max&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Parameter reasoning over statistical evidence — flagship reasoning matters here&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Coach&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;qwen3.7-plus&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Generates the student/parent report in Korean, English, and Chinese from a numeric summary containing no personal data&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;(We also keep &lt;code&gt;qwen3.6-flash&lt;/code&gt; configured for smoke tests and debugging iterations.)&lt;/p&gt;

&lt;p&gt;Splitting roles across models is a real design consideration, not an afterthought: Model Studio grants free quota &lt;strong&gt;per model&lt;/strong&gt;, so separating roles across models separates their quota pools too. Role separation and cost control turned out to be the same decision.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkcfczelcs58f8da4mcc0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkcfczelcs58f8da4mcc0.png" alt=" " width="800" height="314"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  What actually happened on the first Qwen generation
&lt;/h2&gt;

&lt;p&gt;We seeded 18 synthetic LAB-5 sessions engineered so that blank staring would be misclassified as focus under the current parameters, then ran one generation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Total elapsed: 2 minutes 10 seconds.&lt;/strong&gt; (Our earlier CLI-agent rehearsal took 7.)&lt;/p&gt;

&lt;p&gt;Qwen proposed two changes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="nl"&gt;"key"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"blank_stare.stare_dispersion_th"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"from"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.035&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"to"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.055&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
 &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"key"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"blank_stare.max_saccade_count_1s"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"from"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;   &lt;/span&gt;&lt;span class="nl"&gt;"to"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.6&lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Holdout results — five sessions the proposal never saw:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;State&lt;/th&gt;
&lt;th&gt;Sensitivity&lt;/th&gt;
&lt;th&gt;Specificity&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;blank_stare&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.273 → 1.000&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1.000 → 1.000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;focus&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1.000 → 1.000&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.619 → 1.000&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;off_task&lt;/td&gt;
&lt;td&gt;1.000 → 1.000&lt;/td&gt;
&lt;td&gt;1.000 → 1.000&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Zero regressions. Gate passed.&lt;/p&gt;

&lt;p&gt;But the part that made us stop and reread the log was the reasoning attached to the first change:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"…a value distinct from the previously rejected 0.06 / 0.062, to distribute holdout overfitting risk."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Qwen had read the generation history, found that an earlier agent had proposed 0.06 and had it rejected, and picked a different point in the separation gap &lt;strong&gt;specifically to avoid repeating that failure&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Nobody prompted it to do that. It was in the evidence, and it used it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why we call this a memory agent
&lt;/h2&gt;

&lt;p&gt;There's a version of "agent memory" that means storing chat logs and retrieving them later. That's fine, but it isn't what's happening here.&lt;/p&gt;

&lt;p&gt;Our agent's memory is a &lt;strong&gt;versioned lineage of scoring parameters, each attached to the evidence that justified it&lt;/strong&gt;: the mistake log that motivated it, the server-computed before/after numbers, the gate verdict, whether a human adopted it, and whether it was later rolled back.&lt;/p&gt;

&lt;p&gt;That's memory you can audit. When the SFI shown to a parent changes, you can walk backward through generations and see exactly which change caused it, what evidence motivated it, what it measured on data it hadn't seen, and who approved it.&lt;/p&gt;

&lt;p&gt;Experience accumulating as evidence, not as transcript.&lt;/p&gt;

&lt;h2&gt;
  
  
  How this compares to everything else out there
&lt;/h2&gt;

&lt;p&gt;The market is full of "focus apps" — webcam attention trackers, study timers with cameras, classroom monitoring tools. Almost all of them share two properties we consider disqualifying. First, they ship frames or video off the device — exactly what parents and regulators don't want pointed at children. Second, their scoring logic is frozen: fixed heuristics or a black-box model that interprets a 9-year-old and a 40-year-old with the same thresholds, forever.&lt;/p&gt;

&lt;p&gt;Concentration Cube inverts both. Raw video never leaves the desk, and the interpretation itself evolves — against instructor-labeled clinical ground truth, through a holdout gate, with an auditable lineage and a human owning every adoption. We're not aware of another consumer focus product that combines on-device privacy, hospital-protocol ground truth, and a gated self-evolving interpreter. That combination is the moat — and it deepens daily, because every day of data makes the next generation better. As we like to put it: &lt;strong&gt;this device is at its least capable the day you buy it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1j14flmo3nuxvycpfbo1.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1j14flmo3nuxvycpfbo1.jpg" alt=" " width="800" height="266"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the daily evidence comes from
&lt;/h2&gt;

&lt;p&gt;One thing that surprises people: the ground truth feeding this engine is not synthetic-only and not crowd-labeled. Our team runs an &lt;strong&gt;IRB-approved validation program at a general hospital&lt;/strong&gt; — supervised, consented sessions, collected daily. Each session is an instructed experiment: an investigator tells the subject in real time — &lt;em&gt;focus&lt;/em&gt; (actually read), &lt;em&gt;blank-stare&lt;/em&gt; (keep looking at the book, stop reading), &lt;em&gt;look away&lt;/em&gt; (leave the task entirely) — and logs each command at the moment it is spoken, so the 10 Hz eye/pupil stream is synchronized with ground truth to the second. The labeling method is literally called &lt;code&gt;realtime_instructed&lt;/code&gt; in our API. What evolves against those labels is the &lt;strong&gt;interpretation method itself&lt;/strong&gt; — how gaze dispersion, saccades, and pupil dynamics map onto mental states.&lt;/p&gt;

&lt;p&gt;We chose to confront the ethics and safety questions of measuring children's attention through an institutional review process rather than around it. The architectural consequence is that the evolution server has a steady diet: every day of clinical ground truth is fresh evidence, and fresh evidence is a new generation candidate. The product consumers meet is a small cube on a desk. The thing that makes it smarter every day is this loop.&lt;/p&gt;

&lt;p&gt;It also makes the two-server split make sense. The operations server &lt;em&gt;is&lt;/em&gt; the product. The evolution server exists only to grow it — it owns the mistake logs, the lineage, the proposer, and the gates. Growth is a separate system from the thing that grows.&lt;/p&gt;

&lt;h2&gt;
  
  
  The whole stack is Alibaba-native
&lt;/h2&gt;

&lt;p&gt;One thing worth stating plainly: this is an Alibaba-ecosystem project end to end, by choice. The cube itself is sourced for manufacturing through &lt;strong&gt;Accio&lt;/strong&gt;, Alibaba's AI sourcing platform, with PCBA suppliers from the Alibaba network. The backend — both the operations server and the evolution server — runs on &lt;strong&gt;Alibaba Cloud ECS&lt;/strong&gt; in Singapore. And the intelligence is &lt;strong&gt;Qwen on Model Studio&lt;/strong&gt;, over the OpenAI-compatible endpoint: &lt;code&gt;qwen3.7-max&lt;/code&gt; as the proposer that evolves the interpretation, &lt;code&gt;qwen3.7-plus&lt;/code&gt; as the coach that writes trilingual reports, &lt;code&gt;qwen3.6-flash&lt;/code&gt; for smoke tests. Splitting roles across models also splits Model Studio's per-model free quota — role separation and cost control turned out to be the same decision.&lt;/p&gt;

&lt;p&gt;The roadmap only deepens this: age-stratified interpretation means one evolving lineage per 5-year age band — which means more generations, more evaluation, more Qwen. The heart of this product is the evolution server; the engine driving that heart is &lt;strong&gt;Qwen Cloud&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F97c3el4bdtnu4lgzftkz.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F97c3el4bdtnu4lgzftkz.jpg" alt=" " width="624" height="320"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What we'd tell you to steal
&lt;/h2&gt;

&lt;p&gt;If you're building anything where an LLM modifies a system that produces numbers people rely on:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Let the model choose, not report.&lt;/strong&gt; Separate the decision (cheap to verify) from the measurement (expensive to trust). Never let the same entity do both.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Recompute, don't validate prose.&lt;/strong&gt; Don't ask "does this claim look plausible?" Re-run it. A tolerance check against your own execution is worth more than any amount of prompt engineering.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hold out data the proposal can't touch.&lt;/strong&gt; Self-reported improvement on visible data is not improvement.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Put the human at the last step, not the first.&lt;/strong&gt; Review every proposal and you've built a slow human. Review only what survived mechanical verification and a generalization gate, and you've built leverage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Split roles across models.&lt;/strong&gt; Better fit per task, and on Model Studio, separate free-quota pools.&lt;/li&gt;
&lt;/ol&gt;




&lt;p&gt;&lt;em&gt;Built for the Global AI Hackathon Series with Qwen Cloud. Everything runs on Alibaba Cloud — ECS for the stack, Model Studio (Singapore) for the Qwen models.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Code: &lt;a href="https://github.com/S-O-A-TECH/concentration-cube-cloud" rel="noopener noreferrer"&gt;github.com/S-O-A-TECH/concentration-cube-cloud&lt;/a&gt; · Demo video: &lt;a href="https://youtu.be/S3M3wOPD8zE" rel="noopener noreferrer"&gt;youtu.be/S3M3wOPD8zE&lt;/a&gt; · Live demo: &lt;a href="http://47.84.139.15:8200" rel="noopener noreferrer"&gt;http://47.84.139.15:8200&lt;/a&gt; · Track: MemoryAgent&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>qwen</category>
      <category>alibabacloud</category>
      <category>agents</category>
    </item>
  </channel>
</rss>
