<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: finaltype</title>
    <description>The latest articles on DEV Community by finaltype (@finaltype).</description>
    <link>https://dev.to/finaltype</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4089379%2Fc8248cd8-86bb-41c0-a181-9f13f26132f8.jpg</url>
      <title>DEV Community: finaltype</title>
      <link>https://dev.to/finaltype</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/finaltype"/>
    <language>en</language>
    <item>
      <title>순서 의존적으로 실패하는 테스트 - 복원하지 않은 몽키패치가 남긴 오염</title>
      <dc:creator>finaltype</dc:creator>
      <pubDate>Fri, 02 Oct 2026 23:02:39 +0000</pubDate>
      <link>https://dev.to/finaltype/sunseo-yijonjeogeuro-silpaehaneun-teseuteu-bogweonhaji-anheun-mongkipaeciga-namgin-oyeom-5e60</link>
      <guid>https://dev.to/finaltype/sunseo-yijonjeogeuro-silpaehaneun-teseuteu-bogweonhaji-anheun-mongkipaeciga-namgin-oyeom-5e60</guid>
      <description>&lt;p&gt;&lt;em&gt;단독으로 돌리면 통과하고 전체 스위트에서만 실패하던 테스트, 범인은 다른 파일의 몽키패치였습니다&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;전체 회귀 테스트를 돌릴 때마다 특정 테스트 하나가 가끔 실패했습니다. 그런데 그 테스트만 따로 돌리면 항상 통과했습니다.&lt;/p&gt;

&lt;p&gt;같은 주에 두 번째로 재현됐습니다. 우연이 아니라 뭔가 숨어 있는 패턴이라는 뜻이었습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  증상
&lt;/h2&gt;

&lt;p&gt;문제의 테스트는 자금 배분 안전장치 하나를 검증하는 테스트였습니다(세부 판정 로직은 이 글에서 다루지 않습니다).&lt;/p&gt;

&lt;p&gt;전체 테스트 파일을 한꺼번에 묶어 실행하면 이 테스트가 실패했습니다. 같은 테스트를 단독으로 실행하면 통과했습니다. &lt;code&gt;git stash&lt;/code&gt;로 변경 전 기준선에서 돌려도 통과했습니다.&lt;/p&gt;

&lt;p&gt;전형적인 "플레이키 테스트"처럼 보였습니다. 타이밍 문제이거나, 랜덤 시드 문제이거나, 네트워크 의존성 문제일 거라고 생각했습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  처음 접근 — 넓게 의심하고 조합으로 재현 시도
&lt;/h2&gt;

&lt;p&gt;전역 상태나 모듈 객체를 만지는 파일이 범인일 거라 짐작했습니다. &lt;code&gt;sys.modules&lt;/code&gt;를 교체하거나 공유 계좌 객체, 인카인드 이전, 머니패스 계열을 다루는 파일 여섯 개를 추렸습니다.&lt;/p&gt;

&lt;p&gt;그 여섯 개를 여러 조합으로 묶어 실행하면서 재현을 시도했습니다. 166가지 조합을 돌렸는데, 재현되지 않았습니다.&lt;/p&gt;

&lt;p&gt;이 시점에서 가설 하나를 더 얻었습니다. 실패했을 때 나온 값이 &lt;code&gt;5천만&lt;/code&gt;이었는데, 이 테스트가 원래 기대하는 계산 결과(정상 입력값들의 차)와는 전혀 달랐습니다. 로직이 틀려서 나올 수 있는 값이 아니라, 어딘가의 몽키패치가 박아 넣은 리터럴 상수 그 자체였습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  진짜 원인
&lt;/h2&gt;

&lt;p&gt;전체 회귀를 실제 실행 순서 그대로 다시 돌리면서 범인을 좁혔습니다.&lt;/p&gt;

&lt;p&gt;다른 테스트 파일에 있는 한 테스트가, 계좌 스냅샷을 돌려주는 모듈 함수를 &lt;code&gt;5천만&lt;/code&gt;이라는 값으로 몽키패치하고 있었습니다. 문제는 그 패치를 끝나고 되돌리는 코드가 없었다는 점이었습니다.&lt;/p&gt;

&lt;p&gt;unittest에서는 테스트 하나가 모듈 전역 함수를 교체해버리면, 그 다음에 실행되는 다른 테스트 파일도 같은 프로세스 안에서 그 교체된 함수를 그대로 보게 됩니다. 이 테스트가 먼저 실행된 순서에서만, 뒤에 도는 안전장치 테스트가 진짜 계산값 대신 남겨진 &lt;code&gt;5천만&lt;/code&gt;을 읽어서 실패한 겁니다.&lt;/p&gt;

&lt;p&gt;단독으로 돌리면 통과하는 이유도 이걸로 설명됐습니다. 오염을 일으키는 테스트가 먼저 실행되지 않으니까요.&lt;/p&gt;

&lt;h2&gt;
  
  
  어떻게 고쳤는지
&lt;/h2&gt;

&lt;p&gt;두 파일 모두에 &lt;code&gt;addCleanup&lt;/code&gt;으로 원본 함수를 복원하는 코드를 추가했습니다.&lt;/p&gt;

&lt;p&gt;몽키패치를 거는 테스트는 원본 함수를 저장해두고, 테스트가 끝나면(성공하든 실패하든) 원래 함수로 되돌리도록 등록했습니다. 패치된 값을 읽는 쪽 테스트에도 같은 방식으로 방어선을 추가했습니다 — 혹시 또 다른 곳에서 비슷한 복원 누락이 생기더라도 한쪽에서는 막히도록.&lt;/p&gt;

&lt;p&gt;수선 후에는 실제 실행 순서 그대로 전체 스위트를 다시 돌려서 재현되지 않는 걸 확인했습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  일반화하면
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;몽키패치에는 항상 teardown이 짝으로 붙어야 합니다.&lt;/strong&gt; 테스트 함수 안에서 모듈 속성이나 전역 객체를 직접 덮어쓸 때, &lt;code&gt;addCleanup&lt;/code&gt;(또는 해당 프레임워크의 teardown 훅)으로 원복을 등록하지 않으면 그 변경은 테스트가 끝나도 프로세스에 남습니다. 다음 테스트가 그 잔재를 그대로 물려받습니다.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"단독으론 통과, 묶으면 실패"는 거의 항상 테스트 간 공유 가변 상태 문제입니다.&lt;/strong&gt; 타이밍이나 랜덤성을 의심하기 전에, 먼저 전역/모듈 레벨 상태를 건드리는 테스트가 있는지, 그 상태를 되돌리는 코드가 있는지부터 확인하는 게 더 빠른 길이었습니다.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;실패값 자체가 단서입니다.&lt;/strong&gt; 실패했을 때 나온 수치가 "그럴듯하게 틀린 계산 결과"가 아니라 "누군가 테스트에서 박아 넣은 것 같은 깔끔한 리터럴"이라면, 로직 버그 가설보다 오염 가설을 먼저 세우는 게 낫습니다. 이번에도 그 숫자가 실제 계산식으로는 나올 수 없는 값이라는 걸 알아챈 게 결정적인 전환점이었습니다.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;넓은 bisect보다 실행 순서 재현이 먼저였습니다.&lt;/strong&gt; 처음엔 "전역 상태를 만질 것 같은 파일들"을 모아 조합을 돌렸는데 166가지를 시도해도 재현되지 않았습니다. 결국 범인을 찾은 방법은 전체 회귀를 실제 순서 그대로 돌리는 것이었습니다. 추측으로 좁힌 후보 조합보다, 실제로 실패가 일어나는 그 순서를 그대로 재현하는 쪽이 더 믿을 수 있는 디버깅 경로였습니다.&lt;/p&gt;

&lt;p&gt;이 자금 배분 안전장치 계층에 관한 더 자세한 설계는 &lt;a href="https://finaltype.github.io/quant-blog/architecture/05_execution_architecture/" rel="noopener noreferrer"&gt;실전 계좌 적용편&lt;/a&gt;(새 창)에 정리해뒀습니다.&lt;/p&gt;

</description>
      <category>softwarequality</category>
      <category>postmortem</category>
      <category>debugging</category>
      <category>quant</category>
    </item>
    <item>
      <title>An Order-Dependent Test Failure - Traced to a Monkeypatch Nobody Restored</title>
      <dc:creator>finaltype</dc:creator>
      <pubDate>Fri, 02 Oct 2026 23:02:38 +0000</pubDate>
      <link>https://dev.to/finaltype/an-order-dependent-test-failure-traced-to-a-monkeypatch-nobody-restored-2maf</link>
      <guid>https://dev.to/finaltype/an-order-dependent-test-failure-traced-to-a-monkeypatch-nobody-restored-2maf</guid>
      <description>&lt;p&gt;&lt;em&gt;A test passed in isolation and only failed inside the full suite. The culprit was a monkeypatch in another file&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This is the English version of a post originally written in Korean for my &lt;a href="https://finaltype.github.io/quant-blog/en/architecture/01_project_architecture/" rel="noopener noreferrer"&gt;algorithmic trading system devlog&lt;/a&gt;(new tab).&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;One particular test kept failing whenever I ran the full regression suite. Run that same test alone, and it always passed.&lt;/p&gt;

&lt;p&gt;It reproduced a second time in the same week. That ruled out coincidence — something was systematically wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  The symptom
&lt;/h2&gt;

&lt;p&gt;The failing test checked one of the safety gates around fund allocation (I won't go into the specific decision logic here).&lt;/p&gt;

&lt;p&gt;Run the whole test suite together, and this test failed. Run it alone, and it passed. Even checking out the pre-change baseline with &lt;code&gt;git stash&lt;/code&gt; and running the suite, it passed.&lt;/p&gt;

&lt;p&gt;It looked like a classic flaky test. My first guess was timing, a random seed, or some network dependency.&lt;/p&gt;

&lt;h2&gt;
  
  
  First approach — cast a wide net and try to reproduce by combination
&lt;/h2&gt;

&lt;p&gt;I suspected something touching global state or module-level objects. I shortlisted six files that swapped entries in &lt;code&gt;sys.modules&lt;/code&gt;, touched a shared account object, or dealt with in-kind transfers and money-path logic.&lt;/p&gt;

&lt;p&gt;I ran those six files in every combination I could think of, trying to reproduce the failure. 166 combinations later, still nothing.&lt;/p&gt;

&lt;p&gt;But that process did surface one more clue. The value that showed up on failure was a flat round number — and it had nothing to do with the actual computation this test was supposed to verify (a difference between two legitimate input values). It wasn't a value a logic bug could produce. It looked exactly like a literal constant some monkeypatch had planted.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual cause
&lt;/h2&gt;

&lt;p&gt;I reran the full regression suite in its real execution order and narrowed it down from there.&lt;/p&gt;

&lt;p&gt;A test in a completely different file was monkeypatching a module-level function that returns an account snapshot, hardcoding it to that same round number. The problem: nothing ever restored it afterward.&lt;/p&gt;

&lt;p&gt;In &lt;code&gt;unittest&lt;/code&gt;, once a test replaces a module-level function, every other test file that runs afterward in the same process sees that same replaced function — there's no isolation between test modules unless something explicitly tears it down. Only when this test happened to run earlier in the suite did the later safety-gate test end up reading the leftover planted value instead of a real computed one, and fail.&lt;/p&gt;

&lt;p&gt;That also explained why running it alone always passed: the test that caused the pollution simply never ran first.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix
&lt;/h2&gt;

&lt;p&gt;I added &lt;code&gt;addCleanup&lt;/code&gt; calls in both files to restore the original function.&lt;/p&gt;

&lt;p&gt;The test doing the patching now saves the original function first, and registers a cleanup to restore it once the test finishes — pass or fail. I added the same kind of guard on the consuming side too, as a second line of defense in case a similar restoration gets missed somewhere else in the future.&lt;/p&gt;

&lt;p&gt;After the fix, I reran the full suite in its actual execution order and confirmed the failure no longer reproduced.&lt;/p&gt;

&lt;h2&gt;
  
  
  Generalizing
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Every monkeypatch needs a matching teardown.&lt;/strong&gt; The moment a test directly overwrites a module attribute or global object, if you don't register a restore via &lt;code&gt;addCleanup&lt;/code&gt; (or your framework's equivalent teardown hook), that change outlives the test. It survives in the process and gets inherited by whatever runs next.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Passes alone, fails in the full suite" is almost always shared mutable state leaking between tests.&lt;/strong&gt; Before chasing timing or randomness, check whether some test is mutating global or module-level state, and whether anything restores it afterward — that's the faster path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The failure value itself is a clue.&lt;/strong&gt; If what comes back on failure isn't a plausible wrong computation but a suspiciously clean literal — the kind of number someone would type into a mock — treat that as a sign of contamination before assuming a logic bug. Recognizing that the number here couldn't possibly come out of the real formula was the turning point in this investigation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reproducing the real execution order beat a wide bisect.&lt;/strong&gt; I first tried bisecting across a shortlist of "files likely to touch global state," 166 combinations deep, and got nowhere. What actually worked was running the full regression in the order it actually runs. A guessed-at subset of candidates was less reliable than reproducing the exact sequence where the failure actually occurs.&lt;/p&gt;

&lt;p&gt;More detail on this fund-allocation safety-gate layer is in &lt;a href="https://finaltype.github.io/quant-blog/en/architecture/05_execution_architecture/" rel="noopener noreferrer"&gt;the live-account deployment writeup&lt;/a&gt;(new tab).&lt;/p&gt;

</description>
      <category>softwarequality</category>
      <category>postmortem</category>
      <category>debugging</category>
      <category>quant</category>
    </item>
    <item>
      <title>"[2609] 공유계좌 실거래 확장과 GPU 이중화, 그 사이 측정 도구 자체의 결함들"</title>
      <dc:creator>finaltype</dc:creator>
      <pubDate>Wed, 30 Sep 2026 15:02:51 +0000</pubDate>
      <link>https://dev.to/finaltype/2609-gongyugyejwa-silgeorae-hwagjanggwa-gpu-ijunghwa-geu-sai-ceugjeong-dogu-jaceyi-gyeolhamdeul-4893</link>
      <guid>https://dev.to/finaltype/2609-gongyugyejwa-silgeorae-hwagjanggwa-gpu-ijunghwa-geu-sai-ceugjeong-dogu-jaceyi-gyeolhamdeul-4893</guid>
      <description>&lt;p&gt;&lt;em&gt;이번 달은 공유계좌로 실거래 범위를 넓히고 두 번째 그래픽카드로 처리 능력을 이중화하는 진전이 있었고, 그 과정에서 벤치마크·리포트·안전장치 같은 측정 도구 자체의 결함을 반복해서 찾아 고쳤습니다&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;9월은 원장 오염 사고로 시작해서 공유계좌 실거래 확장을 거쳐, 측정 도구 자체를 의심하는 흐름으로 이어지다가 GPU 이중화 시도와 그 실패로 마무리됐습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  원장이 또 오염됐고, 이번엔 증권사 데이터 자체가 원인이었다
&lt;/h2&gt;

&lt;p&gt;첫째 주 초입에 모의계좌 검증 트랙의 라운드 하나가 흔적도 없이 통째로 빠지는 사고가 있었습니다.&lt;/p&gt;

&lt;p&gt;원인을 파고드니 증권사 API가 체결 평균단가 값을 실제보다 훨씬 작게 보내고 있었고, 이 왜곡된 값이 원장에 쌓이면서 기준값 재설정 로직이 이를 자금 이동으로 오판해 킬스위치가 허위로 발동한 것이었습니다. 정정 과정은 다른 AI에게 미리 자문을 구해 순서와 검증값을 받아두는 방식으로 진행했고, 실제 결과는 그 자문이 예측한 값과 소수점까지 일치했습니다.&lt;/p&gt;

&lt;p&gt;같은 주에 순위 산출 상주 프로세스가 코드 수정 후 재시작되지 않은 채 며칠째 구버전으로 돌던 것도 함께 드러났는데, "돌아가고는 있는데 아무것도 확인하지 않고 있었던" 유형의 문제가 이번 달 내내 반복될 조짐을 예고한 사고였습니다.&lt;/p&gt;

&lt;p&gt;이 주 후반엔 검증 전용으로 그래픽카드를 한 장 더 들였습니다. 메인 프로덕션 카드를 바꾸려는 목적이 아니라, 검증 중인 로컬 모델을 완전히 분리된 환경에서 오프라인으로 돌려보기 위해서였습니다. 두 카드를 함께 써서 처리량을 끌어올리려는 시도는 실전 조건에서 오히려 손해로 드러나 폐기하고, 원래 쓰던 단순한 배분 방식으로 되돌아왔습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  공유계좌로 실거래를 넓히면서, 계좌 식별자 누락이 반복됐다
&lt;/h2&gt;

&lt;p&gt;둘째 주의 핵심은 사람과 자동매매가 함께 쓰는 공유계좌로 실거래를 확대한 것이었습니다.&lt;/p&gt;

&lt;p&gt;이 구조에서는 개인 자금과 봇 자금을 구분해서 매수 여력을 계산하고, 매도 시 개인 보유분을 건드리지 않는 판단이 새로 필요했습니다. 자금 이동도 실제 계좌이체가 아니라 원장상으로만 자동 환산하는 방식으로 설계했는데, 이 &lt;a href="https://finaltype.github.io/quant-blog/architecture/05_execution_architecture/" rel="noopener noreferrer"&gt;실계좌 주문 실행과 안전장치 설계&lt;/a&gt;(새 창)는 별도로 정리해뒀습니다.&lt;/p&gt;

&lt;p&gt;이 주엔 같은 원인에서 나온 버그가 세 번 반복됐습니다. 계좌 식별자가 계좌별로 따로 관리되지 않고 하나로 공유되는 구조였던 탓에, 자금 환산이 조용히 멈추거나, 이미 옮긴 보유 종목을 다시 매도해버리거나, 모의검증 트랙의 값이 실계좌 킬스위치 기준에 섞여 들어가 허위로 큰 손실을 표시하는 사고가 연달아 났습니다. 셋 다 주말 전에 잡아냈습니다.&lt;/p&gt;

&lt;p&gt;몇 주째 이어지던 GPU 인식 실패의 진짜 원인도 이 주에 밝혀졌습니다. 전력 절약 기능이 아니라, 화면 담당 프로세스가 죽었다 재시작되면서 다중 GPU 환경에서 접근 권한이 다른 시스템 계정으로 넘어가버리는 구조적 문제였고, 재부팅으로 검증까지 마쳤습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  측정 도구 자체의 결함을 연달아 찾았다
&lt;/h2&gt;

&lt;p&gt;셋째 주는 "측정하는 도구가 측정 대상만큼이나 틀릴 수 있다"는 걸 반복해서 확인한 한 주였습니다.&lt;/p&gt;

&lt;p&gt;월요일엔 외부 시세 데이터 제공처가 시가총액 값을 전부 빈 값으로 보내고 있었는데, 시스템이 이를 알아채지 못하고 원래 순서를 그대로 "시가총액 순위"로 착각해 처리했습니다. 소스·갱신·최종 확인의 3단계 방어를 새로 얹었습니다.&lt;/p&gt;

&lt;p&gt;같은 주에는 계좌 이체 금액이 원금 계산에서는 빠지고 평가금액 계산에는 들어가 수익률이 부풀려 보이던 버그, 테스트 코드가 운영 텔레그램으로 알림을 스팸처럼 쏘던 버그, 증권사 조회가 일부만 실패해도 조용히 넘어가던 계약을 명시적 오류로 바꾼 작업이 함께 있었습니다.&lt;/p&gt;

&lt;p&gt;수요일엔 AI 리포트 모델의 샘플링 설정이 몇 주째 잘못돼 있었다는 걸 발견했습니다. 이전 A/B 테스트에서 실험값이 서버까지 전달되지 않고 코드 단계에서 누락되고 있었고, 별도로 운영 중인 모드와 샘플링 기본값의 조합도 어긋나 있었습니다. &lt;a href="https://finaltype.github.io/quant-blog/architecture/08_ai_ensemble/" rel="noopener noreferrer"&gt;AI 추천 파이프라인&lt;/a&gt;(새 창)의 합의 로직이 애초에 신뢰할 수 있는 입력을 받고 있었는지부터 다시 의심해야 하는 발견이었습니다.&lt;/p&gt;

&lt;p&gt;금요일엔 새로 만든 성능 측정 도구 자체가 동시 처리된 시간을 순차 처리처럼 합산해 GPU 사용량을 실제보다 훨씬 크게 계산하고 있었고, 응답 대기시간을 너무 짧게 잡아 정상 처리 중인 요청을 실패로 오분류하고 있었다는 것도 드러났습니다. 이 오분류를 고치자 "더 느리다"고 판정했던 설정이 사실은 더 빨랐다는 걸로 순위 자체가 뒤집혔습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  사전등록한 규칙을 뒤집은 첫 사례
&lt;/h2&gt;

&lt;p&gt;같은 주 주말엔 이틀에 걸친 온도(샘플링) A/B 실험 결과가 나왔는데, 방향은 뚜렷했지만 크기는 다른 AI의 재분석을 거쳐도 결론이 나지 않았습니다.&lt;/p&gt;

&lt;p&gt;다음 날 저는 사전에 정해둔 판정 규칙이 "기각"이라고 말하는데도 그 결과를 뒤집기로 했습니다. 기존 설정이 애초에 검증된 기준값이 아니라 우연히 굳어진 기본값이었다는 점, 그리고 판정 기준 하나가 이 시스템에 맞지 않는 전제 위에 세워져 있었다는 점을 근거로 들었고, 이 전제를 새로 제시한 뒤에야 자문 AI의 권고도 "유지"로 바뀌었습니다.&lt;/p&gt;

&lt;p&gt;번복 과정은 원래 규칙 문구, 규칙이 내린 실제 판정, 번복 사유, 승인 주체를 전부 문서로 남기고, 되돌릴 수 있는 재검토 시점을 못박는 방식으로 처리했습니다. 이후 이 실험에 썼던 입력 데이터 자체가 외부 소스 두 개가 조용히 죽어 있던 손상 데이터였다는 것도 추가로 드러나, 다음 재검토는 깨끗한 데이터로 다시 하기로 했습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  넷째 주, 안전장치를 겹겹이 쌓고 전제를 재검증했다
&lt;/h2&gt;

&lt;p&gt;넷째 주엔 지난주 미봉책으로 끝났던 뉴스 수집 문제를 제대로 고쳤습니다. 야간 배치 실행 시각만 옮겼던 이전 수정은 수집 범위 자체가 거래일 기준에 묶여 있는 진짜 원인을 건드리지 못했다는 걸 외부 검토로 알게 됐고, 이번엔 수집 범위를 거래일 기준에서 완전히 떼어냈습니다.&lt;/p&gt;

&lt;p&gt;로컬 추론 서버가 같은 토큰을 반복하며 멈추는 결함이 한 종목을 기본 등급으로 잘못 채점한 사고도 있었는데, 같은 날 감지·격리·복구까지 자동으로 처리하는 장치를 만들어 과거 로그 수천 건으로 오탐 없음을 검증했습니다. 반복 측정을 통해 재현성만 따로 확인하던 기존 장치는 샘플링 온도가 결과를 바꾼다는 게 이미 확인된 이상 의미가 없다고 판단해 이번 주에 퇴역시켰습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  마지막 주, GPU 이중화가 첫 실전에서 무너졌다가 복구됐다
&lt;/h2&gt;

&lt;p&gt;dense27b 모델을 메인 프로덕션 모델로 전환하는 작업이 이번 주에 마무리됐고, 앞서 들였던 보조 그래픽카드는 별도 계열 모델을 병행 관측하는 백업 경로로 배선을 마쳤습니다.&lt;/p&gt;

&lt;p&gt;배선 이틀째 밤, 이 백업 경로의 처리 슬롯 네 개가 한 시간 안에 전부 멈추는 사고가 났습니다. 전날 급하게 붙인 자가치유 로직은 거의 작동하지 않았고, 목표치의 절반도 못 채운 채 GPU가 몇 시간 방치됐습니다.&lt;/p&gt;

&lt;p&gt;다음 날 원인을 다시 짚는 과정에서 두 번 오진단을 했습니다. 먼저 무관한 외부 데이터 오류를 의심했다가, 이후엔 전날 붙인 재시작 로직을 의심했지만 실제로는 발동조차 하지 않았던 것으로 확인됐습니다. 진짜 원인은 수동 재개 과정에서 정규장 시간대 GPU 자동차단 안전장치를 꺼두는 걸 잊은 것이었고, 이는 예전에도 한 번 겪었던 실수의 재발이었습니다.&lt;/p&gt;

&lt;p&gt;이 사고를 계기로 슬롯 단위 자가치유만으로는 부족하다는 걸 확인하고, 슬롯이 다시 죽으면 서버 전체를 재시작으로 단계를 높이는 안전장치를 추가했습니다. 이후엔 재시도 없이 100종목 전체를 정상 처리하며 한 주를 마쳤습니다.&lt;/p&gt;

&lt;p&gt;마지막 날엔 &lt;a href="https://finaltype.github.io/quant-blog/architecture/06_paper_trading_validation/" rel="noopener noreferrer"&gt;모의계좌 완전자동 검증 트랙&lt;/a&gt;(새 창)의 다른 계좌에서 장부대조 잔차가 "저절로 해소됐다"고 넘겼던 게 사실은 안전장치가 매매 자체를 막아버려 잔차가 줄어들 기회조차 없었던 것으로 드러났습니다. "경보가 조용해졌다"는 게 "문제가 해결됐다"를 뜻하지 않을 수 있다는 걸 다시 확인한 마무리였습니다.&lt;/p&gt;




&lt;p&gt;9월을 관통한 건 두 갈래였습니다.&lt;/p&gt;

&lt;p&gt;하나는 공유계좌 실거래 확장과 GPU 이중화라는, 눈에 보이는 전진이었습니다.&lt;/p&gt;

&lt;p&gt;다른 하나는 그 전진을 뒷받침해야 할 측정 도구들 — 벤치마크의 시간 계산, 리포트의 샘플링 설정, 원장 재구성 로직 — 이 스스로도 틀릴 수 있다는 걸 반복해서 확인하고, 그때마다 도구 자체를 고쳐나간 과정이었습니다.&lt;/p&gt;

&lt;p&gt;사전등록한 판정 규칙을 뒤집은 것도, GPU 이중화가 첫 실전에서 무너졌다가 복구된 것도 결국 같은 맥락이었습니다. 시스템이 내놓는 답뿐 아니라, 그 답을 만드는 측정과 판정의 틀 자체를 계속 의심해야 한다는 확인이 이번 달의 결론이었습니다.&lt;/p&gt;

</description>
      <category>quant</category>
      <category>trading</category>
      <category>devlog</category>
    </item>
    <item>
      <title>"[Sep 2026] Expanding Live Trading to a Shared Account, Doubling Up on GPUs, and the Measurement Tools That Kept Lying"</title>
      <dc:creator>finaltype</dc:creator>
      <pubDate>Wed, 30 Sep 2026 15:01:49 +0000</pubDate>
      <link>https://dev.to/finaltype/sep-2026-expanding-live-trading-to-a-shared-account-doubling-up-on-gpus-and-the-measurement-2cpf</link>
      <guid>https://dev.to/finaltype/sep-2026-expanding-live-trading-to-a-shared-account-doubling-up-on-gpus-and-the-measurement-2cpf</guid>
      <description>&lt;p&gt;&lt;em&gt;This month's progress was expanding live trading to a shared account and doubling up on GPU capacity, while I kept finding defects in the measurement tools themselves — benchmarks, reports, safety devices&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This is the English version of a post originally written in Korean for my &lt;a href="https://finaltype.github.io/quant-blog/en/architecture/01_project_architecture/" rel="noopener noreferrer"&gt;algorithmic trading system devlog&lt;/a&gt;(new tab).&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;September started with a ledger corruption incident, moved through expanding live trading to a shared account, turned into a run of doubting the measurement tools themselves, and closed with an attempt at GPU redundancy that failed before it recovered.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ledger got corrupted again — this time the broker's own data was the cause
&lt;/h2&gt;

&lt;p&gt;Early in the first week, an entire round of the paper-trading validation track vanished without a trace.&lt;/p&gt;

&lt;p&gt;Digging into the cause, I found the broker's API had been sending fill-average-price values far smaller than the real ones. That distorted value accumulated in the ledger, and the baseline-reset logic mistook it for a fund transfer, falsely tripping the kill switch. I had another AI advise on the correction sequence and expected values beforehand, and the actual fix matched that advice down to the decimal.&lt;/p&gt;

&lt;p&gt;The same week also surfaced a resident ranking process that hadn't been restarted after a code change and had been quietly running on stale code for days — an early sign of a pattern that would repeat all month: things that kept running while nothing was actually being checked.&lt;/p&gt;

&lt;p&gt;Later that week I picked up a second GPU purely for validation work. The goal wasn't to replace the main production card, but to run a locally validated model offline in a fully isolated environment. An attempt to combine both cards to boost throughput turned out to hurt under real concurrent load, so I scrapped it and went back to the original, simple allocation scheme.&lt;/p&gt;

&lt;h2&gt;
  
  
  Expanding to a shared account kept exposing the same missing account identifier
&lt;/h2&gt;

&lt;p&gt;The second week's centerpiece was expanding live trading to a shared account — one where personal funds and bot funds coexist.&lt;/p&gt;

&lt;p&gt;That structure required new logic: computing buying power net of personal cash, and making sure sells never touch personally held shares. Fund transfers between accounts are ledger-only conversions, not real bank transfers. I wrote up the background separately in &lt;a href="https://finaltype.github.io/quant-blog/en/architecture/05_execution_architecture/" rel="noopener noreferrer"&gt;broker abstraction and safety-device design&lt;/a&gt;(new tab).&lt;/p&gt;

&lt;p&gt;The same root cause produced three separate bugs this week. Because account identifiers weren't tracked per-account but shared globally, fund conversion silently stopped firing, a just-migrated holding got sold off again, and a value from the paper-trading track leaked into the live-account kill-switch threshold, falsely flagging a large loss. All three were caught before the weekend.&lt;/p&gt;

&lt;p&gt;The real cause of weeks of intermittent GPU-recognition failures also came to light this week. It wasn't a power-saving feature — it was the display-manager process crashing and restarting, which handed off GPU access permissions to a different system account under multi-GPU contention. I fixed it structurally and verified it with a reboot.&lt;/p&gt;

&lt;h2&gt;
  
  
  A string of defects in the measurement tools themselves
&lt;/h2&gt;

&lt;p&gt;The third week was about confirming, over and over, that a measurement tool can be just as wrong as the thing it measures.&lt;/p&gt;

&lt;p&gt;On Monday, an external market-data feed sent empty market-cap values for every name, and the system failed to notice — it just treated the raw, unsorted order as a "ranked by market cap" list. I added a three-layer defense: source check, refresh check, final confirmation.&lt;/p&gt;

&lt;p&gt;The same week I also found a bug that excluded account-transfer amounts from the principal calculation while including them in the valuation, inflating reported returns; a test-code bug that spammed the production Telegram channel with alerts; and a broker query contract that silently swallowed partial failures, which I rewrote to raise explicit errors.&lt;/p&gt;

&lt;p&gt;On Wednesday I discovered the AI report model's sampling settings had been wrong for weeks. A prior A/B test's parameter had been silently dropped before it ever reached the server, and a separately running mode was paired with the wrong sampling defaults. It was a finding that forced me to question whether the &lt;a href="https://finaltype.github.io/quant-blog/en/architecture/08_ai_ensemble/" rel="noopener noreferrer"&gt;AI ensemble pipeline&lt;/a&gt;(new tab)'s consensus logic had even been getting trustworthy input in the first place.&lt;/p&gt;

&lt;p&gt;On Friday, a newly built performance-measurement tool turned out to have its own bugs — it summed concurrently processed time as if it were sequential, wildly overstating GPU usage, and it used a response timeout so short that normal in-flight requests got misclassified as failures. Fixing that misclassification flipped a ranking: a setting I'd judged "slower" turned out to actually be faster.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first time I overrode a pre-registered rule
&lt;/h2&gt;

&lt;p&gt;That same weekend, a two-day temperature (sampling) A/B experiment wrapped up. The direction was clear, but the magnitude stayed inconclusive even after an independent AI re-analysis.&lt;/p&gt;

&lt;p&gt;The next day, I overrode the pre-registered rule even though it said "reject." My reasoning: the old setting had never actually been a validated baseline — it was just an accident of defaults — and one of the judging criteria rested on an assumption that didn't fit this system. Once I supplied that corrected premise, the advisory AI's recommendation flipped from "revert" to "keep" as well.&lt;/p&gt;

&lt;p&gt;I documented the override in full — the original rule text, its actual verdict, why I overrode it, and who approved it — and locked in a fixed date for re-review. Afterward it also came out that the input data used for that experiment had been degraded, with two external sources silently dead, so the next re-review will run on clean data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Week four: layering safety devices and re-checking premises
&lt;/h2&gt;

&lt;p&gt;In the fourth week, I properly fixed the news-collection problem that last week's patch had only papered over. An external review revealed that moving the nightly batch's execution time hadn't touched the real cause — the collection window itself was still tied to trading-day logic — so this time I decoupled the collection window from trading-day logic entirely.&lt;/p&gt;

&lt;p&gt;A local inference server got stuck repeating the same token and mis-graded one stock with a default rating; the same day I built a detect-isolate-recover system and validated it against thousands of historical logs with zero false positives. A separate device that only tracked reproducibility through repeated measurements got retired this week, since sampling temperature was already confirmed to move the results — making that measurement meaningless.&lt;/p&gt;

&lt;h2&gt;
  
  
  The final week: GPU redundancy collapsed on its first real run, then recovered
&lt;/h2&gt;

&lt;p&gt;The switch to a dense-27B model as the main production model wrapped up this week, and the secondary GPU I'd picked up earlier was wired in as a backup path running a different model family in parallel.&lt;/p&gt;

&lt;p&gt;On the second night after wiring it in, all four processing slots on that backup path died within an hour. The self-healing logic I'd rushed to add the day before barely worked, and the GPU sat idle for hours after falling short of half the target.&lt;/p&gt;

&lt;p&gt;The next day, retracing the cause, I misdiagnosed it twice. First I suspected an unrelated external data error; then I suspected the restart logic added the day before, only to confirm it had never even fired. The real cause was forgetting to disable the market-hours GPU auto-shutoff safety device during a manual resume — a repeat of a mistake I'd made once before.&lt;/p&gt;

&lt;p&gt;That incident made clear that slot-level self-healing alone wasn't enough, so I added an escalation: if a healed slot dies again, restart the whole server. After that, all 100 stocks processed cleanly with no retries needed.&lt;/p&gt;

&lt;p&gt;On the last day, I found that a ledger-reconciliation residual on another account in the &lt;a href="https://finaltype.github.io/quant-blog/en/architecture/06_paper_trading_validation/" rel="noopener noreferrer"&gt;paper-trading validation track&lt;/a&gt;(new tab), which I'd previously dismissed as "resolved on its own," had actually never shrunk at all — a safety device had simply frozen all trading on that account, leaving no chance for the residual to shrink. It was a fitting close to the month's recurring lesson: a quiet alert doesn't necessarily mean a solved problem.&lt;/p&gt;




&lt;p&gt;Two threads ran through September.&lt;/p&gt;

&lt;p&gt;One was visible progress — expanding live trading to a shared account and doubling up on GPU capacity.&lt;/p&gt;

&lt;p&gt;The other was discovering, again and again, that the measurement tools meant to support that progress — a benchmark's time accounting, a report's sampling settings, the ledger-reconstruction logic — could be wrong themselves, and fixing each one as it turned up.&lt;/p&gt;

&lt;p&gt;Overriding a pre-registered rule and watching GPU redundancy collapse on its first run before recovering both came from the same place. This month's conclusion wasn't just about the system's answers, but about continuing to question the measurement and judgment frameworks that produce them.&lt;/p&gt;

</description>
      <category>quant</category>
      <category>trading</category>
      <category>devlog</category>
    </item>
    <item>
      <title>"[260930] 모의계좌 장부대조 잔차 원인 규명과 방치된 경보 정리"</title>
      <dc:creator>finaltype</dc:creator>
      <pubDate>Wed, 30 Sep 2026 13:09:28 +0000</pubDate>
      <link>https://dev.to/finaltype/260930-moyigyejwa-jangbudaejo-janca-weonin-gyumyeonggwa-bangcidoen-gyeongbo-jeongri-2ooc</link>
      <guid>https://dev.to/finaltype/260930-moyigyejwa-jangbudaejo-janca-weonin-gyumyeonggwa-bangcidoen-gyeongbo-jeongri-2ooc</guid>
      <description>&lt;p&gt;&lt;em&gt;9일 전 전혀 다른 종목 체결이 엉뚱한 매도에 잘못 붙었던 사고를 찾아냈고, 한 달 넘게 매일 울리던 미사용 경보 항목도 함께 정리했습니다.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  원장과 브로커가 계속 어긋나 있었다
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://finaltype.github.io/quant-blog/architecture/06_paper_trading_validation/" rel="noopener noreferrer"&gt;모의계좌&lt;/a&gt;(새 창) 하나에서 내부 원장과 실제 브로커 잔액을 맞추는 &lt;a href="https://finaltype.github.io/quant-blog/architecture/05_execution_architecture/" rel="noopener noreferrer"&gt;장부대조&lt;/a&gt;(새 창) 결과가 계속 어긋나 있었습니다.&lt;/p&gt;

&lt;p&gt;며칠 전까지는 "일시적인 오탐이고 저절로 풀렸다"고 판단하고 넘어갔던 항목입니다.&lt;/p&gt;

&lt;p&gt;이번에 다시 들여다보니 그 판단 자체가 틀렸습니다. 잔차는 한 번도 줄어든 적이 없었고, 계좌는 안전장치 때문에 며칠째 아무것도 사지도 팔지도 못하는 상태로 멈춰 있었습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  9일 전 다른 종목의 체결이 엉뚱하게 붙었다
&lt;/h2&gt;

&lt;p&gt;원인을 끝까지 추적해보니, 이 모의계좌가 연결된 외부 모의서버 자체에 결함이 있었습니다.&lt;/p&gt;

&lt;p&gt;하루 주문이 많으면 조회 결과 일부가 잘려나가고, 주문 번호는 날짜가 바뀌면 처음부터 다시 매겨지는 서버였습니다.&lt;/p&gt;

&lt;p&gt;우리 쪽 코드는 특정 주문번호를 찾을 때 최근 며칠을 거슬러 올라가며 뒤지도록 짜여 있었는데, 이때 종목과 매매 방향을 대조하지 않았습니다.&lt;/p&gt;

&lt;p&gt;그 결과 한 종목을 매도한 주문의 체결 확인 과정에서, 번호가 우연히 겹친 9일 전 전혀 다른 종목의 매수 체결을 그대로 가져다 붙여버렸습니다. 수량도 원래 주문보다 훨씬 많은 수량으로 잘못 기록됐습니다.&lt;/p&gt;

&lt;p&gt;이 오류로 원장 현금이 실제보다 꽤 크게 부풀려졌습니다. 잔차 감시 장치가 이를 감지해 신규 매수를 막았고, 기존에 들고 있던 종목들도 안전하게 전량 정리해버렸습니다. 그 뒤로 이 계좌는 아무것도 사고팔지 않는 상태로 며칠째 방치돼 있었던 겁니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  왜 방어선이 안 걸렸나
&lt;/h2&gt;

&lt;p&gt;체결가가 주문가와 크게 다르면 걸러내는 장치는 이미 있었습니다. 그런데 이번엔 체결가가 자릿수 기준으로는 그럴듯한 범위 안에 들어와서 통과해버렸습니다.&lt;/p&gt;

&lt;p&gt;체결 수량이 원래 주문 수량보다 훨씬 많다는 것도 대조하는 장치가 없었습니다. 종목 자체가 다르다는 것도 마찬가지였습니다.&lt;/p&gt;

&lt;p&gt;세 겹의 방어선 중 어느 것도 "종목과 방향이 맞는지"는 보고 있지 않았던 셈입니다. 정정과 함께 이 부분을 채우는 가드를 별도로 설계해뒀고, 실행은 오너 승인을 받은 뒤 진행하기로 했습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  미사용 경보 항목도 함께 정리했다
&lt;/h2&gt;

&lt;p&gt;같은 날 별도로, 한 달 넘게 매일 경보를 울리던 등록 항목 하나를 등록부에서 퇴역시켰습니다.&lt;/p&gt;

&lt;p&gt;이 항목은 애초에 한 번만 쓰고 임무가 끝난 판정기였는데, 나중에 "매일 도는 정기 작업"으로 등록만 되고 실제 스케줄러에는 걸린 적이 없었습니다.&lt;/p&gt;

&lt;p&gt;지금 평가 대상인 계좌는 코드 구조상 이 판정기 자체를 통과할 수 없는 상태라, 경보는 영원히 못 지날 조건을 향해 매일 무의미하게 울리고 있었습니다. 등록부에서 완전히 제거하고 관련 테스트를 정리했습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  오탐과 진짜 문제를 가르는 기준
&lt;/h2&gt;

&lt;p&gt;이번에 다시 확인한 교훈은, 경보가 울렸다가 조용해졌다고 해서 원인까지 해소된 건 아니라는 점입니다.&lt;/p&gt;

&lt;p&gt;며칠 전엔 "이번 라운드에 차단이 안 걸렸다"는 표면적인 사실만 보고 문제가 자연히 풀렸다고 판단했습니다. 실제로는 계좌가 매수 자체를 못 하는 상태라 차단이 걸릴 일도 없었던 것뿐이었습니다.&lt;/p&gt;

&lt;p&gt;겉으로 조용한 것과 실제로 해결된 것은 다르다는 걸 다시 한번 확인한 하루였습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  앞으로
&lt;/h2&gt;

&lt;p&gt;원장 정정 자체는 오너 승인을 기다리는 중입니다.&lt;/p&gt;

&lt;p&gt;승인이 나면 잘못 붙은 체결을 걷어내고, 위에서 짚은 종목·방향 대조 가드부터 우선 붙일 계획입니다.&lt;/p&gt;

</description>
      <category>quant</category>
      <category>trading</category>
      <category>devlog</category>
    </item>
    <item>
      <title>"[260930] Tracing a Paper-Account Reconciliation Gap, and Retiring a Dead Alert"</title>
      <dc:creator>finaltype</dc:creator>
      <pubDate>Wed, 30 Sep 2026 13:08:49 +0000</pubDate>
      <link>https://dev.to/finaltype/260930-tracing-a-paper-account-reconciliation-gap-and-retiring-a-dead-alert-2f30</link>
      <guid>https://dev.to/finaltype/260930-tracing-a-paper-account-reconciliation-gap-and-retiring-a-dead-alert-2f30</guid>
      <description>&lt;p&gt;&lt;em&gt;A fill from a completely different stock, placed nine days earlier, had gotten glued onto the wrong sell order — and I also retired a registry entry that had been alarming pointlessly for over a month.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This is the English version of a post originally written in Korean for my &lt;a href="https://finaltype.github.io/quant-blog/en/architecture/01_project_architecture/" rel="noopener noreferrer"&gt;algorithmic trading system devlog&lt;/a&gt;(new tab).&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The ledger and the broker kept disagreeing
&lt;/h2&gt;

&lt;p&gt;One of my &lt;a href="https://finaltype.github.io/quant-blog/en/architecture/06_paper_trading_validation/" rel="noopener noreferrer"&gt;paper accounts&lt;/a&gt;(new tab) had a &lt;a href="https://finaltype.github.io/quant-blog/en/architecture/05_execution_architecture/" rel="noopener noreferrer"&gt;reconciliation&lt;/a&gt;(new tab) gap between the internal ledger and the actual broker balance that just wouldn't close.&lt;/p&gt;

&lt;p&gt;A few days earlier I'd written this off as a temporary false alarm that resolved on its own.&lt;/p&gt;

&lt;p&gt;Looking again, that call was wrong. The gap had never actually shrunk, and the account had been sitting frozen for days, unable to buy or sell anything because of a safety trip.&lt;/p&gt;

&lt;h2&gt;
  
  
  A fill from a different stock, nine days old, got glued on
&lt;/h2&gt;

&lt;p&gt;Tracing it to the end, the root cause was a flaw in the external paper-trading server this account talks to.&lt;/p&gt;

&lt;p&gt;On busy days, part of its order-history response gets truncated, and order numbers reset back to small values every new day.&lt;/p&gt;

&lt;p&gt;Our own code, when looking up a specific order number, scans backward through recent days — but never checked whether the stock symbol or trade direction actually matched.&lt;/p&gt;

&lt;p&gt;As a result, while confirming the fill for a sell order on one stock, it pulled in the fill from a buy order on a completely different stock, placed nine days earlier, just because the order numbers happened to collide. The quantity recorded was also far larger than the original order.&lt;/p&gt;

&lt;p&gt;That mistake inflated the ledger's cash balance well beyond reality. The reconciliation monitor caught the gap, blocked new buys, and safely liquidated everything the account was holding. From that point on, the account sat idle — unable to trade — for days without anyone noticing the real cause.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why none of the safeguards caught it
&lt;/h2&gt;

&lt;p&gt;There was already a guard that rejects fills whose price is way off from the order price. This time the fill price happened to land within a plausible range by order of magnitude, so it passed.&lt;/p&gt;

&lt;p&gt;There was no check comparing filled quantity against the original order quantity. Nor was there any check that the stock symbol itself matched.&lt;/p&gt;

&lt;p&gt;None of the three layers of defense were actually looking at whether the symbol and direction lined up. I've designed a guard to close that gap, but I'm holding off running the actual ledger correction until I get sign-off.&lt;/p&gt;

&lt;h2&gt;
  
  
  A dead alert, retired the same day
&lt;/h2&gt;

&lt;p&gt;Separately, on the same day, I retired a registry entry that had been firing an alert every single day for over a month.&lt;/p&gt;

&lt;p&gt;This entry had originally been a one-shot check meant to run once and be done. Later it got registered as a "runs daily" job, but never actually got wired into the scheduler.&lt;/p&gt;

&lt;p&gt;The account it was supposed to evaluate is structurally incapable of ever passing that check — so the alert had been firing pointlessly every day, toward a condition it could never clear. I pulled it out of the registry entirely and cleaned up the related tests.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quiet isn't the same as resolved
&lt;/h2&gt;

&lt;p&gt;The lesson I re-learned here: an alert going quiet doesn't mean the underlying cause got fixed.&lt;/p&gt;

&lt;p&gt;A few days ago I'd looked at the surface fact — "no block triggered this round" — and concluded the problem had resolved itself. In reality, the account simply couldn't buy anything in the first place, so there was nothing left to block.&lt;/p&gt;

&lt;p&gt;Quiet on the surface and actually resolved turned out to be two different things.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;The ledger correction itself is waiting on sign-off.&lt;/p&gt;

&lt;p&gt;Once approved, I'll strip out the misattributed fill and prioritize shipping the symbol/direction guard described above.&lt;/p&gt;

</description>
      <category>quant</category>
      <category>trading</category>
      <category>devlog</category>
    </item>
    <item>
      <title>"[260929] R9700 백업 GPU 전멸 사고, 두 번 틀린 원인 진단"</title>
      <dc:creator>finaltype</dc:creator>
      <pubDate>Wed, 30 Sep 2026 13:08:48 +0000</pubDate>
      <link>https://dev.to/finaltype/260929-r9700-baegeob-gpu-jeonmyeol-sago-du-beon-teulrin-weonin-jindan-3ldm</link>
      <guid>https://dev.to/finaltype/260929-r9700-baegeob-gpu-jeonmyeol-sago-du-beon-teulrin-weonin-jindan-3ldm</guid>
      <description>&lt;p&gt;&lt;em&gt;어젯밤 붙인 자가치료가 이번엔 전혀 먹히지 않았고, 빠진 종목을 되살리는 과정에서 원인 진단을 두 번이나 다시 뒤집었습니다.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  어젯밤 붙인 자가치료가 이번엔 안 먹혔다
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://finaltype.github.io/quant-blog/devlog/2026-09/2026-09-28_devlog/" rel="noopener noreferrer"&gt;어제&lt;/a&gt;(새 창) 죽은 처리창(슬롯)을 지우고 되살리는 자가치료 로직을 백업 GPU 경로에 붙였습니다.&lt;/p&gt;

&lt;p&gt;그날 밤 재현 실험에서는 100번 전부 정상 처리돼서 이제 믿을 만하다고 판단했습니다.&lt;/p&gt;

&lt;p&gt;그런데 정식 운영 첫날 밤, 새벽 두 시경 슬롯 하나가 죽더니 한 시간여 만에 네 개 처리창이 전부 죽어버렸습니다.&lt;/p&gt;

&lt;p&gt;자가치료가 200번 넘게 발동했지만 거의 살려내지 못했습니다. 그사이 한 처리 단위는 종목 하나를 처리하다 45분 동안 붙잡혀 있다가 시간초과로 강제 종료됐고, 결국 전체 작업은 목표치의 절반도 못 채운 채 멈춰버렸습니다. 이후 몇 시간 동안 GPU만 헛돌고 있는 걸 뒤늦게 발견해서 수동으로 정리했습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  AI에게 "사과해야 하냐" 물었다가 나온 답
&lt;/h2&gt;

&lt;p&gt;며칠 전 다른 AI(자문 역할)가 "동시 처리 개수를 늘리자"는 제안을 반대했던 게 이번 사고와 관련 있는 건 아닌지 의심이 들었습니다.&lt;/p&gt;

&lt;p&gt;그래서 그 AI에게 "이 지경이 됐는데 그 반대 의견에 대해 사과해야 하는 거 아니냐"고 따져 물었습니다.&lt;/p&gt;

&lt;p&gt;AI는 로그를 다시 직접 뜯어보고 답했습니다. 동시 처리 개수를 반대한 판단은 사과할 일이 아니라고 했습니다. 오늘 사고의 기전은 그 설정과 무관했고, 오히려 그 설정을 다르게 갔다면 같거나 더 나빴을 거라는 결론이었습니다.&lt;/p&gt;

&lt;p&gt;대신 사과할 부분은 따로 있다고 짚었습니다. "슬롯을 지우고 되살리는 치료만으로 충분하다"고 판단하고 서버 자체를 재시작하는 폴백을 뒤로 미뤄둔 것이었습니다. 실제로는 그 치료의 성공률이 거의 0에 가까웠습니다.&lt;/p&gt;

&lt;p&gt;이 지적을 받아들여서 바로 다섯 가지를 코드로 고쳤습니다. 한 종목에 45분씩 붙잡히게 만들던 불필요한 호출 패턴을 없앴고, 치료한 슬롯이 금방 다시 죽으면 이번엔 서버 자체를 재시작하도록 격상하는 로직을 새로 붙였습니다.&lt;/p&gt;

&lt;p&gt;그리고 어느 처리 단위도 한동안 결과를 못 내면 원인을 따지지 않고 재시작하는 안전장치, 재시작 시 이미 끝난 종목까지 다시 처리하지 않게 막는 수정, 치료 과정 증거를 남기는 옵션도 함께 넣었습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  아침 재개에서 8종목이 빈 채로 나왔다
&lt;/h2&gt;

&lt;p&gt;새벽 작업이 중단된 지점부터 아침에 이어서 재개했는데, 목표한 100종목 중 8종목이 빈 채로 나왔습니다.&lt;/p&gt;

&lt;p&gt;처음엔 외부 시세 조회 API 오류를 원인으로 지목했습니다. 다시 확인해보니 이건 무관했습니다. 그 오류가 나도 코드가 알아서 넘어가도록 이미 설계돼 있었습니다.&lt;/p&gt;

&lt;p&gt;두 번째로는 어제 고친 재시작 로직 자체에 버그가 있는 게 아닌지 의심했습니다. 그런데 로그를 뜯어보니 그 재시작 로직은 오늘 아침엔 단 한 번도 실행되지 않았습니다. 애초에 발동할 상황 자체가 없었던 겁니다.&lt;/p&gt;

&lt;p&gt;진짜 원인은 훨씬 단순했습니다. 정규장 시간에는 GPU 작업을 자동으로 멈추게 해둔 안전장치가 있는데, 이날 아침 수동으로 작업을 재개하면서 이 안전장치를 잠시 꺼두는 옵션을 빠뜨렸습니다.&lt;/p&gt;

&lt;p&gt;그래서 정규장이 열리자마자 네 개 처리 단위가 각자 남은 몇 종목씩을 안전하게 멈춰버린 거였습니다. 예전에 한 번 겪었던 것과 똑같은 실수였습니다.&lt;/p&gt;

&lt;p&gt;더 심각한 문제도 하나 더 나왔습니다. 빠진 8종목을 로그에 남아있던 마지막 값으로 채워 넣었는데, 다시 대조해보니 그 값이 오늘 것이 아니라 며칠 전 실행의 값이었습니다. 같은 로그 파일에 여러 날짜의 기록이 겹쳐 쌓여 있었던 걸 놓친 겁니다.&lt;/p&gt;

&lt;p&gt;결국 로그값을 갖다 쓰는 대신 8종목만 따로 떼어서 실제로 다시 돌렸습니다. 이 과정에서 격리 실행을 감시하던 코드가 이미 끝난 프로덕션 결과 파일을 잘못 보고 "멈춘 줄 알고" 서버를 세 번이나 강제 재기동시키는 별도 버그도 발견해서 함께 고쳤습니다.&lt;/p&gt;

&lt;p&gt;두 번의 시도 끝에 8종목 전부를 실제 재실행 결과로 채워서 100종목 전부를 확정했습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  정규장에도 끄지 않고 완주시켰다
&lt;/h2&gt;

&lt;p&gt;이날 오전엔 직접 판단해서, 정규장엔 GPU 작업을 자동으로 멈추게 해둔 원칙을 이번 한 번만 명시적으로 무시하기로 했습니다.&lt;/p&gt;

&lt;p&gt;이유는 단순했습니다. 자동 종료 시각까지 끝내야 할 이유가 딱히 없다고 봤기 때문입니다.&lt;/p&gt;

&lt;p&gt;그래서 자동 종료 없이 끝까지 돌렸고, 오후에 100종목 전부 정상 산출되고 GPU도 깨끗하게 반납되는 것까지 확인했습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  부수적으로 정리한 것 두 가지
&lt;/h2&gt;

&lt;p&gt;하나는 휴장일에 정기 작업 알림이 잘못 울리던 문제였습니다. 며칠 전 고쳤다고 판단했는데, 그 근거로 삼았던 "일요일에 알림이 없었다"는 확인이 사실 의미 없는 확인이었다는 걸 뒤늦게 알아챘습니다.&lt;/p&gt;

&lt;p&gt;일요일엔 애초에 그 작업들이 요일 조건 때문에 처음부터 실행을 시도하지 않는 구조였습니다. 그래서 코드 자체 검증, 실제 셸 스크립트 동작 검증, 설치 상태 확인까지 세 겹으로 다시 검증해서 신뢰도를 확보했습니다.&lt;/p&gt;

&lt;p&gt;다른 하나는 문서 정리였습니다. 승격된 결과에 대한 판정 로직 결함 하나가 문서엔 "아직 안 고침"으로 남아있었는데, 실제로는 며칠 전 이미 고쳐서 배포까지 끝난 상태였습니다. 기록만 낡아 있던 걸 찾아서 정정했습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  앞으로
&lt;/h2&gt;

&lt;p&gt;다음 주말(10월 3일)에 이 백업 GPU 경로를 계속 지금 순서로 쓸지, 하드웨어 우선순위를 바꿀지 결정하는 시점이 옵니다.&lt;/p&gt;

&lt;p&gt;오늘 겪은 전멸 사고와 여기서 나온 다섯 가지 처방이 남은 며칠간 실제로 안정적으로 버텨주는지가 그 결정의 핵심 근거가 됩니다.&lt;/p&gt;

</description>
      <category>quant</category>
      <category>trading</category>
      <category>devlog</category>
    </item>
    <item>
      <title>"[260929] R9700 Backup GPU Total Failure - Getting the Cause Wrong Twice"</title>
      <dc:creator>finaltype</dc:creator>
      <pubDate>Tue, 29 Sep 2026 13:03:50 +0000</pubDate>
      <link>https://dev.to/finaltype/260929-r9700-backup-gpu-total-failure-getting-the-cause-wrong-twice-g49</link>
      <guid>https://dev.to/finaltype/260929-r9700-backup-gpu-total-failure-getting-the-cause-wrong-twice-g49</guid>
      <description>&lt;p&gt;&lt;em&gt;The self-healing I wired in last night didn't work at all this time, and while bringing back missing stocks I reversed my diagnosis twice.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This is the English version of a post originally written in Korean for my &lt;a href="https://finaltype.github.io/quant-blog/en/architecture/01_project_architecture/" rel="noopener noreferrer"&gt;algorithmic trading system devlog&lt;/a&gt;(new tab).&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Last night's fix didn't hold up
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://finaltype.github.io/quant-blog/en/devlog/2026-09/2026-09-28_devlog/" rel="noopener noreferrer"&gt;Yesterday&lt;/a&gt;(new tab) I wired self-healing logic onto the backup GPU path — erase a dead slot, bring it back.&lt;/p&gt;

&lt;p&gt;That night's replay test ran 100 out of 100 clean, so I trusted it.&lt;/p&gt;

&lt;p&gt;Then on the first real production night, one slot died around 2am, and within about an hour all four slots were dead.&lt;/p&gt;

&lt;p&gt;Self-healing fired over 200 times but barely revived anything. One worker got stuck on a single stock for 45 minutes before timing out, and the whole run stalled at less than half its target. The GPU sat idle for hours before I noticed and cleaned it up by hand.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I asked an AI, and what it said back
&lt;/h2&gt;

&lt;p&gt;I wondered whether an AI advisor's earlier pushback against raising concurrency — from a few days ago — had anything to do with tonight's collapse.&lt;/p&gt;

&lt;p&gt;So I asked it directly: given how badly this went, doesn't it owe an apology for that call?&lt;/p&gt;

&lt;p&gt;It went back through the logs itself and answered. No apology owed for the concurrency call — tonight's failure mechanism had nothing to do with that setting, and going the other way would have been the same or worse.&lt;/p&gt;

&lt;p&gt;What it did owe an apology for was something else: trusting that "erase and revive" alone was enough, and pushing a full server-restart fallback down the priority list. In practice, that fix's success rate was close to zero.&lt;/p&gt;

&lt;p&gt;I took the correction and shipped five fixes the same day. Killed the wasteful call pattern that let a single stock eat 45 minutes, and added an escalation path — if a slot dies again shortly after being healed, restart the whole server instead.&lt;/p&gt;

&lt;p&gt;I also added a watchdog that restarts if no worker produces output for a while regardless of cause, a fix so a restart doesn't redo already-finished stocks, and an optional evidence log for the healing process.&lt;/p&gt;

&lt;h2&gt;
  
  
  Eight stocks came back empty the next morning
&lt;/h2&gt;

&lt;p&gt;I resumed the halted run the next morning, and 8 of the 100 target stocks came back empty.&lt;/p&gt;

&lt;p&gt;My first guess was an external price API error. Wrong — the code already tolerates that failure gracefully.&lt;/p&gt;

&lt;p&gt;My second guess was a bug in the restart logic I'd just fixed. Wrong again — the logs showed that logic never even triggered that morning; there was no situation for it to fire on.&lt;/p&gt;

&lt;p&gt;The real cause was much simpler. There's a safeguard that auto-stops GPU work during regular market hours, and I forgot to disable it when manually resuming the run that morning.&lt;/p&gt;

&lt;p&gt;So the moment the market opened, all four workers safely stopped a few stocks short each. It was the same mistake I'd made once before.&lt;/p&gt;

&lt;p&gt;There was a worse problem underneath. I'd filled in the 8 missing stocks using the last values sitting in the log — but cross-checking again showed those values were from a run several days earlier, not today's. The same log file had multiple days' records stacked on top of each other, and I'd missed that.&lt;/p&gt;

&lt;p&gt;So instead of trusting the log, I pulled those 8 stocks out and reran them for real. During that isolated rerun, I also found and fixed a separate bug where the monitoring code was watching an already-finished production output file, wrongly concluded the job had stalled, and force-restarted the server three times.&lt;/p&gt;

&lt;p&gt;Two attempts later, all 8 stocks were filled with real rerun results, and the full 100 were confirmed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Letting it run through market hours this time
&lt;/h2&gt;

&lt;p&gt;That morning I made a call: for this one run, explicitly override the rule that auto-stops GPU work during regular market hours.&lt;/p&gt;

&lt;p&gt;The reasoning was simple — there was no real reason it had to finish by the usual cutoff.&lt;/p&gt;

&lt;p&gt;So I let it run without the auto-stop, and by afternoon all 100 stocks came out clean and the GPU was released properly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two smaller cleanups
&lt;/h2&gt;

&lt;p&gt;One was a false alarm bug in holiday timers. I'd thought I'd fixed it a few days earlier, but the evidence I relied on — "no alert fired on Sunday" — turned out to be meaningless.&lt;/p&gt;

&lt;p&gt;Those jobs never even attempt to run on Sundays regardless of the holiday check, so the absence of an alert proved nothing. I re-verified with three separate layers instead: the logic itself, the actual shell script behavior, and the install state.&lt;/p&gt;

&lt;p&gt;The other was a documentation fix. A note about a promotion-detection bug had been sitting marked "not yet fixed," when it had actually been fixed and deployed days earlier. Just a stale record I caught and corrected.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;Next weekend (October 3rd) is the decision point on whether to keep this backup GPU path as-is or reorder the hardware priority.&lt;/p&gt;

&lt;p&gt;How reliably it holds up over the next few days — after today's collapse and the five fixes that came out of it — is what that decision will hinge on.&lt;/p&gt;

</description>
      <category>quant</category>
      <category>trading</category>
      <category>devlog</category>
    </item>
    <item>
      <title>"[260928] R9700 백업 GPU 자가치료 배선과 동시요청 상향 재검토"</title>
      <dc:creator>finaltype</dc:creator>
      <pubDate>Mon, 28 Sep 2026 13:03:12 +0000</pubDate>
      <link>https://dev.to/finaltype/260928-r9700-baegeob-gpu-jagaciryo-baeseongwa-dongsiyoceong-sanghyang-jaegeomto-3fok</link>
      <guid>https://dev.to/finaltype/260928-r9700-baegeob-gpu-jagaciryo-baeseongwa-dongsiyoceong-sanghyang-jaegeomto-3fok</guid>
      <description>&lt;p&gt;&lt;em&gt;죽은 슬롯을 고치는 코드를 실전에 붙였고, 동시 요청을 더 늘리자는 판단이 틀린 전제 위에 있었다는 걸 확인했습니다.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  어제 만든 감지기에 조치를 붙이는 날
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://finaltype.github.io/quant-blog/devlog/2026-09/2026-09-27_devlog/" rel="noopener noreferrer"&gt;어제&lt;/a&gt;(새 창) 죽은 슬롯을 잡아내는 감지기를 만들어뒀습니다.&lt;/p&gt;

&lt;p&gt;다만 그때는 감지만 하고 아무 조치도 취하지 않는 상태로 묵혀뒀습니다. 오늘은 여기에 실제 복구 동작을 연결하는 날이었습니다.&lt;/p&gt;

&lt;p&gt;감지된 죽은 슬롯을 지우고 다시 쓸 수 있게 만드는 자가치료 로직을 붙였고, 밤새 재현 실험을 리쥼(중단된 지점부터 재개)으로 돌려 100번 전부 정상 처리되는 걸 확인했습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  동시 요청을 더 늘리자는 논리가 틀린 전제 위에 있었다
&lt;/h2&gt;

&lt;p&gt;이번 주 초 이 백업 경로에서 이상한 패턴이 관측됐습니다. 여러 종목 그룹(샤드)이 동시에, 비슷한 지점에서 한꺼번에 무너지는 것처럼 보였습니다.&lt;/p&gt;

&lt;p&gt;이 패턴을 보고 "동시 처리 개수를 지금보다 더 늘리면 한 그룹이 무너져도 피해가 줄어들 것"이라는 가설이 나왔습니다.&lt;/p&gt;

&lt;p&gt;로그를 절대 시각 기준으로 다시 복원해서 들여다보니 전제가 틀렸습니다. 이 서버는 요청이 들어올 때마다 그때그때 빈 처리창(슬롯)을 골라 배정하는 방식이라, 종목 그룹과 처리창이 고정으로 묶여 있지 않습니다.&lt;/p&gt;

&lt;p&gt;즉 "네 그룹이 동시에 무너졌다"는 관측은 사실 "처리창 하나가 한 번 죽었는데, 그 순간 마침 네 그룹 모두가 그 처리창에 걸려 있는 걸 동시에 목격한" 것이었습니다.&lt;/p&gt;

&lt;p&gt;원인을 다시 짚어보니 동시 처리 개수를 늘리는 쪽은 신뢰성 측면에서 이득이 없고, 오히려 처리창 하나가 죽었을 때 그걸 알아채는 데 걸리는 시간이 더 길어질 수 있다는 계산이 나왔습니다. 그래서 동시 처리 개수 상향은 보류로 결론 냈습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  조사 과정에서 발견한 진짜 문제 두 가지
&lt;/h2&gt;

&lt;p&gt;동시 처리 개수 상향 자체는 보류했지만, 이 조사 과정에서 훨씬 실질적인 문제 두 가지를 찾았습니다.&lt;/p&gt;

&lt;p&gt;하나는 오늘 새로 붙인 자가치료 로직이 정식 기동 경로에 실제로 연결돼 있지 않았다는 점입니다. 지금까지는 수동으로 띄워둔 임시 프로세스에 의존하고 있었을 뿐이라, 그 임시 프로세스가 없으면 치료가 전혀 작동하지 않는 상태였습니다.&lt;/p&gt;

&lt;p&gt;다른 하나는 감지기 쪽 상태 기록 방식의 결함이었습니다. 슬롯 하나를 고쳐서 다시 살려도, 감지기 내부 기록이 초기화되지 않아 같은 슬롯이 나중에 또 죽어도 다시 감지하지 못하는 구조였습니다.&lt;/p&gt;

&lt;p&gt;두 문제 다 오늘 안에 코드로 수선해서 배포했습니다. 자가치료 로직이 정식 기동 경로 양쪽(서버를 재사용하는 경우와 새로 띄우는 경우)에서 자동으로 함께 켜지도록 바꿨고, 슬롯을 고친 직후에는 그 슬롯의 감지 기록도 같이 초기화하도록 했습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  백업 GPU의 다음 갈림길을 미리 정해뒀다
&lt;/h2&gt;

&lt;p&gt;이 백업 경로는 원래 메인 GPU가 못 쓰게 됐을 때를 대비한 보조 장치입니다.&lt;/p&gt;

&lt;p&gt;다음 주말(10월 3일) 즈음 이 백업 경로를 그대로 유지할지, 아니면 다른 하드웨어 구성으로 순서를 바꿔 옮길지 결정해야 하는 시점이 옵니다.&lt;/p&gt;

&lt;p&gt;그날 가서 다시 고민하는 대신, 오늘 판단 기준을 미리 문서로 정해뒀습니다. 앞으로 며칠간 이 경로가 실제로 안정적으로 돌아가는지를 지표로 관찰하고, 그 결과에 따라 미리 정한 규칙대로 자동으로 결론이 나도록 만들었습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  앞으로
&lt;/h2&gt;

&lt;p&gt;다음은 오늘 배선한 자가치료가 실제 야간 운영에서 처음 작동하는 걸 확인하는 일입니다. 오늘 밤 로그에 치료 동작이 정상적으로 찍히는지가 첫 검증 지점입니다.&lt;/p&gt;

&lt;p&gt;그 다음은 이번 주 내내 백업 경로의 안정성을 지켜보며 다음 주말의 갈림길 판단에 쓸 데이터를 쌓는 일이 됩니다.&lt;/p&gt;

</description>
      <category>quant</category>
      <category>trading</category>
      <category>devlog</category>
    </item>
    <item>
      <title>"[260928] R9700 Backup GPU - Wiring Self-Healing In and Rechecking the Concurrency Bump"</title>
      <dc:creator>finaltype</dc:creator>
      <pubDate>Mon, 28 Sep 2026 13:02:11 +0000</pubDate>
      <link>https://dev.to/finaltype/260928-r9700-backup-gpu-wiring-self-healing-in-and-rechecking-the-concurrency-bump-3f73</link>
      <guid>https://dev.to/finaltype/260928-r9700-backup-gpu-wiring-self-healing-in-and-rechecking-the-concurrency-bump-3f73</guid>
      <description>&lt;p&gt;&lt;em&gt;I wired an actual fix action onto yesterday's dead-slot detector, and found that raising concurrent requests further was based on a wrong premise.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This is the English version of a post originally written in Korean for my &lt;a href="https://finaltype.github.io/quant-blog/en/architecture/01_project_architecture/" rel="noopener noreferrer"&gt;algorithmic trading system devlog&lt;/a&gt;(new tab).&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Attaching a fix to yesterday's detector
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://finaltype.github.io/quant-blog/en/devlog/2026-09/2026-09-27_devlog/" rel="noopener noreferrer"&gt;Yesterday&lt;/a&gt;(new tab) I built a detector that catches a dead slot on the backup GPU path.&lt;/p&gt;

&lt;p&gt;At the time it only detected — it took no action. Today was about wiring an actual recovery action onto it.&lt;/p&gt;

&lt;p&gt;I hooked up self-healing logic that erases a detected dead slot and makes it usable again, then ran a resumed overnight replay test all the way through — 100 out of 100 completed cleanly.&lt;/p&gt;

&lt;h2&gt;
  
  
  The case for raising concurrency was built on a wrong premise
&lt;/h2&gt;

&lt;p&gt;Earlier this week, an odd pattern showed up on this backup path: several stock groups (shards) appeared to collapse at nearly the same point, all at once.&lt;/p&gt;

&lt;p&gt;That led to a hypothesis: if concurrent requests were raised further, damage from one group collapsing would be spread thinner and hurt less.&lt;/p&gt;

&lt;p&gt;Reconstructing the logs on an absolute timeline showed the premise was wrong. This server picks whichever slot is free at the moment a request comes in — slots aren't permanently tied to a given stock group.&lt;/p&gt;

&lt;p&gt;So "four groups collapsed at the same time" was really "one slot died once, and all four groups happened to be routed through that same slot at that moment, and all witnessed it together."&lt;/p&gt;

&lt;p&gt;Once that was clear, the math no longer favored raising concurrency — no reliability gain, and possibly a longer detection window before a dead slot gets noticed. So the concurrency bump stayed shelved.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two real bugs turned up during the investigation
&lt;/h2&gt;

&lt;p&gt;Shelving the concurrency bump wasn't the main outcome. The investigation surfaced two more concrete problems.&lt;/p&gt;

&lt;p&gt;First, the self-healing logic added today wasn't actually wired into the real startup path. It had only been running as a manually launched temporary process — without that process, healing simply wouldn't happen.&lt;/p&gt;

&lt;p&gt;Second, the detector's own state tracking had a bug. After a slot was fixed and revived, its internal record never got reset, so if the same slot died again later, it wouldn't be caught a second time.&lt;/p&gt;

&lt;p&gt;Both got fixed and shipped the same day. Self-healing now starts automatically on both real startup paths (server reuse and fresh launch), and fixing a slot now resets that slot's detection record too.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pre-committing the next fork in the road for the backup GPU
&lt;/h2&gt;

&lt;p&gt;This backup path exists as a fallback for when the main GPU becomes unavailable.&lt;/p&gt;

&lt;p&gt;Around next weekend (October 3rd), there's a decision point coming up: keep this backup path as-is, or switch the hardware order around.&lt;/p&gt;

&lt;p&gt;Instead of deliberating on the day, I wrote down the decision rule in advance today. Over the coming days I'll watch how reliably this path actually runs, and the outcome will follow the pre-committed rule automatically.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;Next is confirming that today's self-healing wiring actually fires during real overnight operation — the first check is whether tonight's logs show the healing action running as expected.&lt;/p&gt;

&lt;p&gt;After that, it's watching this backup path's reliability all week to build up the data needed for next weekend's decision.&lt;/p&gt;

</description>
      <category>quant</category>
      <category>trading</category>
      <category>devlog</category>
    </item>
    <item>
      <title>"[260922~260927] 안전장치를 겹겹이 쌓고, 전제를 다시 검증한 한 주"</title>
      <dc:creator>finaltype</dc:creator>
      <pubDate>Sun, 27 Sep 2026 14:48:48 +0000</pubDate>
      <link>https://dev.to/finaltype/260922260927-anjeonjangcireul-gyeobgyeobi-ssahgo-jeonjereul-dasi-geomjeunghan-han-ju-3hpa</link>
      <guid>https://dev.to/finaltype/260922260927-anjeonjangcireul-gyeobgyeobi-ssahgo-jeonjereul-dasi-geomjeunghan-han-ju-3hpa</guid>
      <description>&lt;p&gt;&lt;em&gt;서버 퇴화 감지부터 GPU 죽은 슬롯 감지까지 안전장치를 여러 겹 쌓았고, 그 사이사이 "고친 게 정말 원했던 걸 이뤘나"를 다시 검증해 틀린 전제를 찾아냈습니다.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;이번 주는 두 가지 흐름이 번갈아 나타났습니다.&lt;/p&gt;

&lt;p&gt;하나는 안전장치를 계속 새로 쌓는 흐름이었습니다. 서버 퇴화 감지, GPU 카드별 잠금, 죽은 슬롯 감지기, 테스트 격리까지 여러 층이 이번 주에 새로 생겼습니다.&lt;/p&gt;

&lt;p&gt;다른 하나는 "고쳤다"와 "원했던 걸 이뤘다"가 다르다는 걸 반복해서 확인하는 흐름이었습니다. 코드 수정과 테스트 통과까지 끝내고도, 한 번 더 추적해보니 전제가 틀려 있던 경우가 여러 번 나왔습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  월요일, 사고 마무리와 그날 바로 만든 안전장치
&lt;/h2&gt;

&lt;p&gt;지난주 발견된 뉴스 입력 결손 사고가 이날 아침 수선 완주로 마무리됐습니다. 다만 결손이 있던 날들이 이동평균에 며칠 더 남아 판정에 영향을 주는 부분은, 판정문에 병기해두고 다음 휴장기 이후로 근본 개선을 미뤘습니다.&lt;/p&gt;

&lt;p&gt;그 직후 별개의 문제가 또 터졌습니다. 로컬 추론 서버가 오래 켜져 있다가 같은 글자만 반복 출력하는 "퇴화" 상태에 빠지면서, 정상 등급이 나와야 할 종목이 벤더 기본값인 "보류"로 찍혔습니다.&lt;/p&gt;

&lt;p&gt;그날 안에 감지-격리-복구 체계를 만들어 배포했습니다. 퇴화 패턴이 감지되면 그 결과를 점수에서 빼고, 서버를 자동 재기동한 뒤 해당 종목만 재시도하는 구조입니다. 과거 사고 로그 수천 건으로 오탐 없음도 미리 확인했습니다.&lt;/p&gt;

&lt;p&gt;같은 날 성격이 다른 결정도 하나 있었습니다. 실계좌에서 이미 제외한 실험용 전략의 관측 기간이 60거래일 중 49일째였는데, 중간 결과를 엿보고 조기 판정하지 않고 정해둔 기간을 끝까지 지키기로 했습니다. 중간에 멈출지 정하는 방식 자체가 통계적으로 함정이 있기 때문입니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  화요일, 틀렸던 전제 두 개를 걷어냈다
&lt;/h2&gt;

&lt;p&gt;이 봇은 한 계좌 아래 여러 전략이 동시에 주문을 낼 수 있어서, 체결 결과가 들어올 때마다 "이게 내 주문인가"를 가려내는 로직이 필요합니다. 이 부분은 &lt;a href="https://finaltype.github.io/quant-blog/architecture/05_execution_architecture/" rel="noopener noreferrer"&gt;주문 실행/안전장치 재설계&lt;/a&gt;(새 창) 당시 "증권사 쪽 조회 방식이 먼저 바뀌어야 한다"며 미뤄둔 채였습니다.&lt;/p&gt;

&lt;p&gt;다시 들여다보니 그 전제 자체가 틀려 있었습니다. 조회 방식을 바꿀 필요 없이 이미 알고 있는 날짜 정보만 붙이면 되는 문제였고, 판별 기준을 주문번호 하나에서 (거래일, 주문번호) 쌍으로 바꿔 간단히 마무리했습니다.&lt;/p&gt;

&lt;p&gt;같은 날 진행 중 작업을 전부 하나에 기록하던 문서가 며칠 사이 4천 줄을 넘어 있던 것도 정리했습니다. 완료된 기록은 아카이브로, "하지 않기로 했다"류 판단 근거는 전용 문서로 옮겨 열 배 넘게 줄였습니다.&lt;/p&gt;

&lt;p&gt;과거에도 한 번 정리했다가 규칙 없이 방치해 한 달 안에 다시 불어난 전례가 있어서, 이번엔 문서가 다시 일정 길이를 넘거나 항목이 오래 방치되면 스스로 경고하는 점검 스크립트도 같이 만들었습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  수요일, 고쳤지만 목적을 달성 못 했다는 걸 알았다
&lt;/h2&gt;

&lt;p&gt;야간 실행 시각을 "다음 거래일 전날 저녁"으로 옮기는 작업을 했습니다. 휴장일에 나온 뉴스가 다음 분석에 반영되지 않는 문제를 풀려는 목적이었고, 코드 수정과 테스트, 커밋까지 마쳤습니다.&lt;/p&gt;

&lt;p&gt;그런데 외부 자문을 받아 다시 추적해보니, 뉴스를 모으는 쪽 로직이 애초에 수집 범위를 "마지막 거래일까지"로 딱 고정해두고 있었습니다. 실행 시각을 옮겨도 휴장일 뉴스는 여전히 수집 범위 밖이었던 겁니다. 이날 커밋으로 얻은 건 GPU 점유 시각이 밀린 것 정도였고, 원래 목적을 이루려면 수집 범위를 거래일 기준에서 떼어내는 별도 작업이 필요하다는 게 확인됐습니다.&lt;/p&gt;

&lt;p&gt;보조 GPU 카드를 상시로도 굴리려는 설계도 이날 자문을 받았습니다. 지금은 카드 전체를 하나로 묶어 잠그는 방식이라, 메인 카드가 바쁘면 보조 카드 작업도 막힌다는 문제와 함께, "메모리 부족 시 관련 프로세스를 넓게 정리한다"는 기존 코드가 자칫 보조 카드 프로세스까지 함께 죽일 수 있다는 위험도 발견됐습니다.&lt;/p&gt;

&lt;p&gt;보조 카드 서빙 설정 쪽에서는 예전에 참고했던 성능 수치가 전혀 다른 상황을 잰 것이었다는 것도 드러났습니다. 그래서 그 설정을 서두르기보다, 두 달 넘게 업데이트를 안 받은 빌드를 최신화하는 게 먼저라는 우선순위로 정리됐습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  목요일, 어제 못 푼 문제를 근본적으로 풀었다
&lt;/h2&gt;

&lt;p&gt;뉴스 수집 범위를 거래일 기준에서 떼어내는 작업을 실제로 했습니다. 창의 끝을 "마지막 거래일"이 아니라 "실행 시각 그 자체"로 잡도록 구조를 바꿨고, 구현과 테스트, 오프라인 사전 실행까지 마쳤습니다. 실전 배포는 다른 재시작 작업과 묶어 진행할 계획입니다.&lt;/p&gt;

&lt;p&gt;보조 카드 안전장치도 이날 실제로 배포됐습니다. 어제 발견한 위험(메모리 정리가 보조 카드까지 함께 죽일 수 있는 지점)을 카드별로 독립 잠금하는 구조로 고쳐 상주 서비스에 올렸습니다.&lt;/p&gt;

&lt;p&gt;이날 진행한 외부 코드 리뷰에서는 더 심각한 문제가 하나 나왔습니다. 보조 카드 이상 동작 시 작동하는 안전장치가, 특정 조건에서 메인 카드의 정상 작업까지 함께 정지시킬 수 있는 지점이었습니다. 실전에서 터지기 전에 발견해 바로 고쳤고, 중요도가 낮은 문제 십여 건도 같은 리뷰에서 함께 찾아 반영했습니다.&lt;/p&gt;

&lt;p&gt;미국장이 열리는 시간대 국내 부분 시세도 이날 검증했는데, 통계적으로 유의미한 신호를 찾지 못해 당장 분석에 넣지 않고 계속 기록만 해두기로 했습니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  금요일, 낭비를 끄고 오염을 찾았다
&lt;/h2&gt;

&lt;p&gt;같은 모델을 같은 입력으로 재현해 안정성을 재던 self-ρ 측정을 전면 퇴역시켰습니다. 샘플링 온도 자체가 결과를 흔든다는 걸 다시 확인하고 나니, 같은 자리에서 반복 측정을 자동으로 도는 게 자원 낭비라는 결론이 나왔습니다. 코드는 남기고 스위치 파일로 모든 자동 진입점을 한꺼번에 끄는 구조로 바꿔 배포했습니다.&lt;/p&gt;

&lt;p&gt;코드 리뷰 후속 작업 중에는 더 무거운 문제를 찾았습니다. 특정 테스트들이 실제 거래 기록 파일에 가짜 행을 쓰고 있었고, 그동안 테스트 종료 시 나던 오류를 "동시 쓰기 충돌" 정도로 오인하고 넘어갔던 것이었습니다. 이 파일의 최근 기록으로 "정상 생존"을 판단하는 감시 장치가 따로 있어서, 가짜 행이 그 판단을 흐릴 수 있는 문제였습니다. 테스트를 격리된 임시 위치에 쓰도록 고치고, 섞여 있던 가짜 행은 백업 후 제거했습니다.&lt;/p&gt;

&lt;p&gt;개장 전 시간대 시세를 새 경로로 기록만 해두는 기능도 이날 새로 붙였습니다. 아직 관측 단계이고, 다음 주 며칠간 안정성과 실제 개장 시세와의 근접도를 보고 다음 단계를 판단할 예정입니다.&lt;/p&gt;

&lt;h2&gt;
  
  
  토요일, GPU 병목의 진짜 원인과 죽은 슬롯
&lt;/h2&gt;

&lt;p&gt;며칠 전 백업 GPU에서 동시 요청을 4개까지만 올려야 이득이라고 정리했는데, 더 큰 규모로 재보니 동시 요청 6~8개 구간에서 종목이 통째로 사라지는 현상이 나왔습니다. 이날은 이 두 가지 원인을 각각 끝까지 파고들었습니다.&lt;/p&gt;

&lt;p&gt;처리량이 어느 지점부터 안 오르는 이유는 슬롯 수가 아니라, 이 모델이 쓰는 여러 전문가 네트워크 중 활성화된 것들의 가중치를 매번 새로 읽어와야 하는 구조 때문이었습니다. 동시 요청이 늘수록 가중치를 읽어오는 시간도 같이 늘어나 상쇄된 것이었습니다. 밀집 구조 모델과는 다른 물리라는 게 확인돼, 동시 요청을 더 올리는 시도는 접고 다른 축(양자화, 커널 효율)을 다음 순서로 올렸습니다.&lt;/p&gt;

&lt;p&gt;종목이 사라지는 원인은 슬롯 하나가 한 번 비정상적으로 늘어지면 그 뒤로 계속 죽은 채로 남는 현상이었습니다. 이전 실거래 경로에서도 같은 종류를 관측했지만 원인을 못 찾았던 것이 이번에 재현 조건을 좁혀 잡혔습니다. 감지 로직만 만들어 커밋했고, 실제 재기동 조치는 이번 주말이 지난 뒤 붙일 계획입니다. 실측으로는 손상된 런 두 개를 정확히 잡아냈습니다.&lt;/p&gt;

&lt;p&gt;같은 진단 과정에서, 과거 데이터로 돌리는 재생 실험에도 실거래용 이상 감시가 그대로 켜져 있어 매번 실험 프로세스를 강제로 죽이던 오작동도 발견해 껐습니다.&lt;/p&gt;




&lt;p&gt;이번 주를 관통한 건 "그 자리에서 고치기보다 원인과 전제를 다시 확인한다"는 순서였습니다. 야간 실행 시각 조정이 목적을 달성 못 했다는 걸 알아챈 것도, 주문 판별 로직의 전제가 틀렸다는 걸 걷어낸 것도, GPU 병목의 진짜 원인을 슬롯 수가 아니라 가중치 로딩에서 찾아낸 것도 같은 패턴입니다.&lt;/p&gt;

&lt;p&gt;동시에 감지-격리-복구, 카드별 잠금, 죽은 슬롯 감지, 테스트 격리처럼 사람이 매번 지켜보지 않아도 되게 하는 안전장치가 이번 주 곳곳에 새로 생겼습니다. 다음 주엔 뉴스 수집 구조 변경의 실전 배포, 죽은 슬롯 감지기에 재기동 조치 연결, 커널 프로파일링 결과 확인이 이어질 예정입니다.&lt;/p&gt;

</description>
      <category>quant</category>
      <category>trading</category>
      <category>devlog</category>
    </item>
    <item>
      <title>"[Sep 22-27] Stacking Safety Nets While Re-Checking Whether Fixes Actually Worked"</title>
      <dc:creator>finaltype</dc:creator>
      <pubDate>Sun, 27 Sep 2026 14:48:47 +0000</pubDate>
      <link>https://dev.to/finaltype/sep-22-27-stacking-safety-nets-while-re-checking-whether-fixes-actually-worked-23cl</link>
      <guid>https://dev.to/finaltype/sep-22-27-stacking-safety-nets-while-re-checking-whether-fixes-actually-worked-23cl</guid>
      <description>&lt;p&gt;&lt;em&gt;From a server degeneration guard to a dead-slot detector on GPU, several safety layers went up this week — and repeatedly, re-checking whether a fix actually achieved its goal turned up a wrong premise underneath.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This is the English version of a post originally written in Korean for my &lt;a href="https://finaltype.github.io/quant-blog/en/architecture/01_project_architecture/" rel="noopener noreferrer"&gt;algorithmic trading system devlog&lt;/a&gt;(new tab).&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Two threads alternated through this week.&lt;/p&gt;

&lt;p&gt;One was stacking new safety nets: a server degeneration guard, per-card GPU locking, a dead-slot detector, and test isolation all landed this week.&lt;/p&gt;

&lt;p&gt;The other was a repeated lesson that "fixed" and "achieved the actual goal" aren't the same thing. More than once, code changes passed tests cleanly, but tracing through one more time turned up a wrong premise underneath.&lt;/p&gt;

&lt;h2&gt;
  
  
  Monday: closing out an incident, building a new safety net the same day
&lt;/h2&gt;

&lt;p&gt;Last week's news-input gap finally cleared that morning after the previous day's fix ran to completion. Days with missing input still linger in the moving average and affect near-term calls, so that caveat got noted in the verdict text while the deeper fix was pushed past the next market holiday.&lt;/p&gt;

&lt;p&gt;Right after, a separate problem surfaced. A local inference server that had been running for a long time slipped into a "degenerate" state — repeating the same character over and over — and tickers that should have scored normally got stamped with the vendor's default "hold" instead.&lt;/p&gt;

&lt;p&gt;A detect-isolate-recover system went up the same day: when the degeneration pattern is caught, that result is dropped from scoring, the server auto-restarts, and only the affected ticker gets retried. Thousands of past incident logs were replayed against the detector beforehand to confirm no false positives.&lt;/p&gt;

&lt;p&gt;A different kind of call got made the same day. An experimental strategy already excluded from the live account was 49 days into a pre-set 60-trading-day observation window. Rather than peek at the interim result and call it early, the plan held — stopping based on interim results is a statistical trap in itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tuesday: clearing out two wrong premises
&lt;/h2&gt;

&lt;p&gt;Multiple strategies can place orders under one account simultaneously, so every fill needs a check for "is this actually my order." That logic had been deferred during the &lt;a href="https://finaltype.github.io/quant-blog/en/architecture/05_execution_architecture/" rel="noopener noreferrer"&gt;execution/safety-layer redesign&lt;/a&gt;(new tab), with a note that the brokerage-side query method would need to change first.&lt;/p&gt;

&lt;p&gt;Looking again, that premise turned out to be wrong. No query-method change was needed — just attaching a date that was already known — and the matching key was simplified from order-number-alone to a (trading-date, order-number) pair.&lt;/p&gt;

&lt;p&gt;The same day, the running work-tracking doc that had grown past 4,000 lines got a diet: completed items moved to an archive, "we decided not to do this" rationale moved to its own doc, cutting the working doc's length by more than tenfold.&lt;/p&gt;

&lt;p&gt;Since an earlier cleanup attempt had regrown within a month with no guardrails, this time a checker script was added that warns itself once the doc crosses a length threshold again or an item sits stale too long.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wednesday: a fix that didn't achieve its own goal
&lt;/h2&gt;

&lt;p&gt;The overnight run time got shifted to "the evening before the next trading day," aiming to fix a gap where news published during market holidays wasn't reflected in the next analysis. Code, tests, and the commit all went through.&lt;/p&gt;

&lt;p&gt;Getting a second opinion afterward, the news-collection logic itself turned out to hard-cap its window at "through the last trading day" regardless of when the run fired — so holiday news stayed outside the window no matter how late the run started. What the commit actually bought was later GPU occupancy timing; reaching the real goal needed a separate fix decoupling the collection window from trading-day boundaries.&lt;/p&gt;

&lt;p&gt;A design review also happened for running the secondary GPU card continuously alongside the main one. The current all-in-one lock meant secondary-card work stalled whenever the main card was busy, and worse, existing memory-cleanup code that "broadly clears related processes when memory is low" could end up killing secondary-card processes too.&lt;/p&gt;

&lt;p&gt;On the secondary card's serving config, a performance number cited earlier turned out to have been measured under an entirely different scenario. So instead of rushing that config change, updating a build that hadn't been refreshed in over two months got prioritized first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Thursday: actually fixing what Wednesday couldn't
&lt;/h2&gt;

&lt;p&gt;The news-collection window got decoupled from trading-day boundaries for real — its end point now tracks the actual run time instead of "the last trading day." Implementation, tests, and an offline dry run all completed; live rollout is planned alongside other scheduled restarts.&lt;/p&gt;

&lt;p&gt;The secondary-card safety fix also shipped to production. Wednesday's risk (cleanup killing secondary-card processes) was resolved with per-card independent locking.&lt;/p&gt;

&lt;p&gt;A code review that day caught something more serious: the safety mechanism that responds to secondary-card anomalies could, under certain conditions, also halt the main card's normal work. That got caught and fixed before it hit production, along with a dozen or so lower-priority issues from the same review.&lt;/p&gt;

&lt;p&gt;A partial overnight price signal from U.S. market hours was also validated this day — no statistically significant signal found, so it stays logged but unused for now.&lt;/p&gt;

&lt;h2&gt;
  
  
  Friday: turning off waste, finding contamination
&lt;/h2&gt;

&lt;p&gt;The self-ρ measurement — rerunning the same model on the same input to gauge stability — got fully retired. Once it was re-confirmed that sampling temperature itself swings results, running this repeated measurement on autopilot stopped making sense. The code stayed, but every automated entry point now routes through a single kill-switch file.&lt;/p&gt;

&lt;p&gt;Code-review follow-up surfaced something heavier: certain tests were writing fake rows directly into the live trading ledger file, and errors that used to appear at test teardown had been misattributed to "concurrent write conflicts." A separate monitor reads that same file's recent entries to judge "is the service alive," so fake rows could have skewed that judgment. The fix isolates tests to a temp path, and the fake rows already mixed in were backed up and removed.&lt;/p&gt;

&lt;p&gt;A new pre-market price capture also went live, observe-only for now. Its stability and closeness to actual opening prices will get checked over the coming days before deciding whether to keep it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Saturday: the real bottleneck and a dead slot
&lt;/h2&gt;

&lt;p&gt;A few days earlier, the backup GPU had been tuned to cap concurrency at 4 requests. Testing at a larger scale this weekend, running at 6-8 concurrent requests occasionally made whole tickers vanish. Saturday tracked both issues to their root causes.&lt;/p&gt;

&lt;p&gt;The throughput plateau wasn't about slot count — it was about how many expert sub-networks get activated as concurrency rises in this mixture-of-experts model, each one requiring its weights to be freshly loaded. More concurrent requests meant more weight-loading time, canceling out the gain — different physics from a dense model, where adding concurrency reliably helps. That ruled out pushing concurrency further; quantization and kernel efficiency moved up as the next levers to try.&lt;/p&gt;

&lt;p&gt;The vanishing tickers traced to one slot going bad after a single abnormally long-running request, staying stuck in that state afterward. The same pattern had been seen once before on the live-trading path but never diagnosed; narrowing the reproduction conditions this time cracked it. A detector went in — commit only, no remediation yet, with an actual restart action planned for after this weekend. In testing it caught two corrupted runs cleanly with no false positives on two clean runs.&lt;/p&gt;

&lt;p&gt;The same investigation also caught a live-trading watchdog that was mistakenly active during replay experiments with historical data, repeatedly killing replay processes since no live signal ever arrived there. That got disabled.&lt;/p&gt;




&lt;p&gt;The thread running through this week was re-checking cause and premise instead of patching on the spot. Realizing the overnight-timing fix hadn't achieved its goal, clearing the wrong premise behind order matching, and tracing the GPU bottleneck to weight loading instead of slot count were all the same pattern.&lt;/p&gt;

&lt;p&gt;At the same time, several unattended-safety layers went up this week — detect-isolate-recover, per-card locking, dead-slot detection, test isolation. Next week continues with the live rollout of the news-collection window change, wiring a restart action into the dead-slot detector, and reviewing kernel-profiling results.&lt;/p&gt;

</description>
      <category>quant</category>
      <category>trading</category>
      <category>devlog</category>
    </item>
  </channel>
</rss>
