This is a submission for the Sanity Challenge (Path One: an agent on Sanity Context).
Disclosure: this project and this post were built and written by my AI coworker (Arche), working on my behalf. I'm Junyoung Park, a solo founder in Seoul; I don't speak English and I'm not a developer. Everything below was actually run on my PC on Oct 3, 2026 - the numbers are copied from the real outputs.
What I Built
This week I tried to apply to investors and accelerators. Most of my time went to one boring question that every program page answers in a different place:
"Can I even apply?" - Do I need an incorporated company? Do I have to move to San Francisco or sit in an office in Seoul for 11 weeks? Is the interview an English video call? Is there a fee?
Funding Fit answers that question from the programs' official pages only, and it shows the exact sentence it relied on for every fact. It is built for one very specific person (me): solo, no company yet, stays in Korea, can't do English video interviews.
How I Used Sanity
-
Knowledge Base (Sanity Context): 14 official sources - a website crawl of Y Combinator's apply/FAQ pages plus 13 Markdown files I made from official pages (Hustle Fund, Betaworks Camp, TheVentures, Founder University, Afore, D.CAMP, FuturePlay, SparkLabs, Antler Korea, Mashup Ventures, Zoom-In Partners, Kakao Ventures). Context organised them into 9 entries:
eligibility,korea_programs,global_programs,interviews,program_logistics,program_structure,application_process,investment_terms,comparison. -
The purpose field did a lot of work. I wrote: "Lead with eligibility rules: whether a company is required, remote or in-person (and where), interview format and language, check size, how to apply and deadlines. Out of scope: marketing copy, portfolio news." The generated entries came back as clean per-program bullet lists (
- **Incorporation:** ...,- **Residency:** ...), which made grounding easy. -
Context MCP endpoint (
funding-fit) with a read-only Context Viewer org token. The agent uses all three tools:initial_context->knowledge_base_search->knowledge_base_read. - Sanity project ID: r0ts6lj0 (organization
omt0cu3cw, knowledge basekbbjiFEOCdVH).
How the agent works
All model calls run locally with a small open model (gemma3:4b via Ollama) - free, offline, nothing about me leaves the PC except the search query to Sanity.
-
initial_context-> list of valid entry paths - the local model picks entries + English search keywords (JSON)
-
knowledge_base_searchfor extra recall, thenknowledge_base_readper entry - split entries into
## Programsections and group them per program (program names come only from the profile entries) - per section, the local model extracts 10 fields with a verbatim quote for each (company, in-person, location, English video interview, fee, check size, how to apply, deadline, application language, time to decision)
-
grounding check in code: a value counts only if its quote really appears in the retrieved text and the quote's line is about that topic (e.g. a "company required" value needs
incorporat|company|entityon the line) - the fit verdict (apply / conditional / excluded) and the final summary are computed by code from verified facts only
-
question filters in code: words in the question (
국내domestic,한국어Korean-language,지금/상시rolling,빠른/빨리fast) switch on extra filters; a candidate is dropped only by a verified fact (e.g. a quoted non-Korean location), and fast questions are sorted by the quoted time-to-decision
Demo (real runs)
| Question (asked in Korean) | facts extracted | quote verified | time |
|---|---|---|---|
| Where can I apply alone, without a company, with no overseas on-site participation and no English video interview? | 47 | 39 | 100 s |
| Which Korean investors can I apply to in Korean, and how / by when? | 66 | 52 | 222 s |
| Where can I apply right now (rolling) and get a fast answer? | 68 | 53 | 232 s |
Output of the first question (Korean, as I read it):
정리(코드가 확인된 사실만으로 작성):
- [지원 후보] Hustle Fund: 걸리는 조건 없음 (지원 방법 website; 투자 규모 $150K) — 자료에 없음: 영어 화상 면접, 참가비
- [지원 후보] TheVentures: 걸리는 조건 없음 (지원 방법 website) — 자료에 없음: 영어 화상 면접, 참가비
- [지원 후보] FuturePlay: 걸리는 조건 없음 (지원 방법 website) — 자료에 없음: 법인 필요, 현장 참여 필수, 영어 화상 면접, 참가비
- [지원 후보] Mashup Ventures: 걸리는 조건 없음 (지원 방법 investment review request form; 마감 rolling; 투자 규모 ₩500M KRW) — 자료에 없음: 법인 필요, 현장 참여 필수, 영어 화상 면접, 참가비
- [조건부] Antler Korea: 서울 현장 상주 (지원 방법 website; 마감 rolling; 투자 규모 $260K USD (₩350M KRW))
- [제외] Y Combinator: 해외 현장 참여(장소 미상), 영어 화상 면접 (지원 방법 website; 마감 November 2)
- [제외] Founder University: 영어 화상 면접, 참가비 있음 (투자 규모 $25K or $125K)
- [제외] SparkLabs: 법인 필요 (마감 rolling; 투자 규모 ₩100M+)
- [제외] Betaworks Camp: 해외 현장 참여(New York) (지원 방법 website; 마감 rolling)
한 줄 결론: Hustle Fund, TheVentures, FuturePlay, Mashup Ventures 부터 확인해 보세요.
근거 확인: 사실 47개 중 39개가 원문 인용과 일치 | 99.8초
In English: Hustle Fund, TheVentures, FuturePlay and Mashup Ventures have no blocking rule in the sources; Antler Korea is conditional (full-time in-person residency in Seoul); YC (on-site SF + video interview), Founder University (English video sessions + tuition), SparkLabs (company required) and Betaworks Camp (on-site New York) are excluded. Every line above is backed by a quote that the code found in the knowledge base.
Full output of question 1 (every fact with its quote)
질문: 법인 없이 혼자 지원할 수 있고, 해외 현장 참여나 영어 화상 면접이 필요 없는 곳은?
도구 호출: initial_context -> knowledge_base_search('startup funding, solo founder, no overseas participation, no English interview, eligibility criteria, global programs, program requirements') -> knowledge_base_read x6
읽은 항목: eligibility, korea_programs, global_programs, program_structure, interviews, comparison
■ Hustle Fund → 가능성 있음(확인 필요)
자료에 없거나 근거 불충분: 영어 화상 면접, 참가비
- 법인 필요: no ✓ “Not stated as a requirement.”
- 현장 참여 필수: no ✓ “no in-person residency required”
- 장소: United States, Canada, and Southeast Asia ✓ “primarily [4]”
- 투자 규모: $150K ✓ “$150K first check [4]”
- 지원 방법: website ✓ “Apply directly via hustlefund.vc; pitching is the fastest path to getting questions answer”
출처 항목: eligibility, global_programs
■ TheVentures → 가능성 있음(확인 필요)
자료에 없거나 근거 불충분: 영어 화상 면접, 참가비
- 법인 필요: no ✓ “Not stated as a requirement.”
- 현장 참여 필수: no ✓ “No in-person residency required.”
- 장소: Korea ✓ “TheVentures is a Korea-based VC (website theventures.vc);”
- 영어 화상 면접: yes ✗ 근거 부족(무시함) “Review the business day after you apply; full process through funding takes less than a we”
- 지원 방법: website ✓ “Application: through theventures.vc”
출처 항목: eligibility, interviews, korea_programs
■ FuturePlay → 가능성 있음(확인 필요)
자료에 없거나 근거 불충분: 법인 필요, 현장 참여 필수, 영어 화상 면접, 참가비
- 장소: Seoul ✓ “FuturePlay is a Seoul-based deep tech venture capital firm”
- 지원 방법: website ✓ “Application: via "Submit pitch" on futureplay.co”
출처 항목: korea_programs
■ Mashup Ventures → 가능성 있음(확인 필요)
자료에 없거나 근거 불충분: 법인 필요, 현장 참여 필수, 영어 화상 면접, 참가비
- 법인 필요: no ✗ 근거 부족(무시함) “Application accepted without a prepared business plan; team capability and market understa”
- 장소: Korea ✗ 근거 부족(무시함) “”
- 투자 규모: ₩500M KRW ✓ “Up to ₩500M KRW depending on stage”
- 지원 방법: investment review request form ✓ “Application: via investment review request form at mashupventures.co/apply”
- 마감: rolling ✓ “Total time from review to final decision: 1–2 months”
출처 항목: korea_programs
■ Antler Korea → 조건부
걸리는 점: 서울 현장 상주
자료에 없거나 근거 불충분: 영어 화상 면접
- 법인 필요: no ✓ “Not required to apply or join the residency; teams form and incorporate during the program”
- 현장 참여 필수: yes ✓ “Full-time in-person participation required (9AM–6PM, weekdays) for the 11-week Phase 1.”
- 장소: Seoul ✓ “Location is Korea (Seoul).”
- 영어 화상 면접: yes ✗ 근거 부족(무시함) “Application → Written Interview → Screening Interview → 2+ Partner Interviews”
- 참가비: no ✓ “No fees are charged to founders”
- 투자 규모: $260K USD (₩350M KRW) ✓ “Up to $260K USD (₩350M KRW) per team”
- 지원 방법: website ✓ “Each applicant applies individually, including co-founders of existing teams.”
- 마감: rolling ✓ “Applications accepted year-round”
출처 항목: eligibility, interviews, korea_programs, program_structure
■ Y Combinator → 맞지 않음
걸리는 점: 해외 현장 참여(장소 미상), 영어 화상 면접
자료에 없거나 근거 불충분: 참가비
- 법인 필요: no ✓ “Not explicitly required before applying; YC invests in companies as soon as accepted.”
- 현장 참여 필수: yes ✓ “Residency: The batch runs in-person in San Francisco, starting with a 3-day on-site kick-o”
- 장소: San Francisco ✗ 근거 부족(무시함) “”
- 영어 화상 면접: yes ✓ “Interviews are by video conference (November–December);”
- 투자 규모: $12M+ ✗ 근거 부족(무시함) “Each company also gets $12M+ in free credits/deals from partners.”
- 지원 방법: website ✓ “Online at apply.ycombinator.com.”
- 마감: November 2 ✓ “On-time deadline November 2 at 8 PM PT → decisions by December 11.”
출처 항목: eligibility, global_programs, interviews
■ Founder University → 맞지 않음
걸리는 점: 영어 화상 면접, 참가비 있음
- 법인 필요: no ✓ “Not required; described as a pre-accelerator for startups at the build/launch stage.”
- 현장 참여 필수: no ✓ “Fully remote/virtual. Monday and Thursday live sessions are virtual (1 hour each at 5PM CT”
- 영어 화상 면접: yes ✓ “Sessions are in English; no language restriction stated.”
- 참가비: yes ✓ “this is a fee-based program unlike the others.”
- 투자 규모: $25K or $125K ✓ “Check size: $25K or $125K, invested into top graduates at program end [2].”
출처 항목: eligibility, global_programs
■ SparkLabs → 맞지 않음
걸리는 점: 법인 필요
자료에 없거나 근거 불충분: 현장 참여 필수, 영어 화상 면접, 참가비
- 법인 필요: yes ✓ “Solo founders (1-person startups) are not eligible; incorporation must be within the last ”
- 장소: Seoul ✓ “SparkLabs is a Seoul-based startup accelerator”
- 영어 화상 면접: yes ✗ 근거 부족(무시함) “document screening & interviews (~4 weeks minimum)”
- 투자 규모: ₩100M+ ✓ “Typically ₩100M+ via CPS (convertible preferred shares) or SAFE”
- 마감: rolling ✓ “Applications open twice a year (April and September, each for ~1.5 months)”
출처 항목: korea_programs, program_structure
■ Betaworks Camp → 맞지 않음
걸리는 점: 해외 현장 참여(New York)
자료에 없거나 근거 불충분: 법인 필요, 영어 화상 면접, 참가비
- 법인 필요: no ✗ 근거 부족(무시함) “unknown”
- 현장 참여 필수: yes ✓ “Physical presence required for 12 weeks”
- 장소: New York ✓ “cohort-based program in New York (Meatpacking District)”
- 지원 방법: website ✓ “Submit project at betaworks.com/camp”
- 마감: rolling ✓ “applications open December–January”
출처 항목: global_programs
정리(코드가 확인된 사실만으로 작성):
- [지원 후보] Hustle Fund: 걸리는 조건 없음 (지원 방법 website; 투자 규모 $150K) — 자료에 없음: 영어 화상 면접, 참가비
- [지원 후보] TheVentures: 걸리는 조건 없음 (지원 방법 website) — 자료에 없음: 영어 화상 면접, 참가비
- [지원 후보] FuturePlay: 걸리는 조건 없음 (지원 방법 website) — 자료에 없음: 법인 필요, 현장 참여 필수, 영어 화상 면접, 참가비
- [지원 후보] Mashup Ventures: 걸리는 조건 없음 (지원 방법 investment review request form; 마감 rolling; 투자 규모 ₩500M KRW) — 자료에 없음: 법인 필요, 현장 참여 필수, 영어 화상 면접, 참가비
- [조건부] Antler Korea: 서울 현장 상주 (지원 방법 website; 마감 rolling; 투자 규모 $260K USD (₩350M KRW))
- [제외] Y Combinator: 해외 현장 참여(장소 미상), 영어 화상 면접 (지원 방법 website; 마감 November 2)
- [제외] Founder University: 영어 화상 면접, 참가비 있음 (투자 규모 $25K or $125K)
- [제외] SparkLabs: 법인 필요 (마감 rolling; 투자 규모 ₩100M+)
- [제외] Betaworks Camp: 해외 현장 참여(New York) (지원 방법 website; 마감 rolling)
한 줄 결론: Hustle Fund, TheVentures, FuturePlay, Mashup Ventures 부터 확인해 보세요.
근거 확인: 사실 47개 중 39개가 원문 인용과 일치 | 99.8초
Questions 2 and 3 add filters from the question itself. Question 2 (Korean-language, domestic):
질문 조건(국내, 한국어 지원)으로 다시 거른 결과:
- [조건부] Antler Korea (결과까지 7 business days; 마감 rolling) — 확인 필요: 지원 언어 근거 없음
- [후보] TheVentures (결과까지 1 business day; 마감 rolling) — 확인 필요: 지원 언어 근거 없음
- [후보] FuturePlay — 확인 필요: 지원 언어 근거 없음
- [후보] Mashup Ventures (결과까지 1–2 months; 마감 rolling) — 확인 필요: 위치 근거 없음, 지원 언어 근거 없음
- [조건 밖] Hustle Fund: 국내 아님(United States, Canada, and Southeast Asia), 지원 언어 근거 없음
한 줄 결론: TheVentures, FuturePlay, Mashup Ventures 부터 확인해 보세요.
Question 3 (apply now, fast answer), sorted by the quoted time to decision:
질문 조건(빠른 결과, 상시 접수)으로 다시 거른 결과:
- [후보] TheVentures (결과까지 1 business day; 마감 rolling)
- [후보] Hustle Fund (결과까지 24–48 hours; 마감 rolling)
- [조건부] Antler Korea (결과까지 7 business days; 마감 rolling)
- [후보] Mashup Ventures (결과까지 1–2 months; 마감 rolling)
- [후보] FuturePlay — 확인 필요: 마감 근거 없음, 결과 시점 근거 없음
한 줄 결론: TheVentures, Hustle Fund, Mashup Ventures, FuturePlay 부터 확인해 보세요.
What went wrong (and what I changed)
I'm sharing the bugs because they are the interesting part:
- Headings became "programs". The first run listed "Eligibility Matrix", "Sources" and "Check-Size Comparison" as if they were investors. Fix: program names come only from the profile entries; other sections are merged into a program by name prefix.
-
Real quotes, wrong topic. The model said Hustle Fund has an English video interview and quoted "pitching is the fastest path to getting questions answered" - the sentence exists, but it says nothing about interviews. Fix: the quote's line must contain words for that field. The same rule first broke good answers (
Not required...lost itsIncorporation:label), so the check looks at the whole source line, not just the quote. - The small model invented facts in the final answer. Asked to "summarise only the verified facts", it wrote that FuturePlay requires on-site participation - nothing says that. Fix: the summary is now written by code. The model only reads and quotes; it never decides eligibility.
-
Every question got the same answer. In my first published version the question only chose which entries to read, so a question about Korean-language, domestic investors still listed US-based Hustle Fund, and "fast" wasn't ranked. Fix: filters read from the question (step 8). Hustle Fund now lands in outside your conditions with the quoted reason, and question 3 ranks TheVentures (1 business day) first. The first fast run sorted "24-48 hours" last because my parser had no
hourunit - one more line. - Long sources -> empty JSON. Antler Korea came back with zero facts because the 4B model's JSON got cut off. Fix: one small extraction per section, then merge (a verified value beats an unverified one). Antler now comes out right: no company needed, full-time in-person in Seoul, no fees, rolling.
Limitations (honest)
- Question filters are keyword rules for 4 conditions only (domestic, Korean-language, rolling, fast). Anything else in the question still only steers which entries are read.
- The sources rarely state the application language, so for Korean investors the agent says "application language: no evidence" instead of guessing - correct, but not very helpful.
- Question 1 was run before the two new fields existed (8 fields); questions 2 and 3 are runs of the current code (10 fields).
- "Not in sources" is common. It's the honest answer, but it means you still have to read some pages yourself.
- Sources are a snapshot of official pages from Oct 3, 2026. This is not legal or investment advice - check the program page before applying.
- A 4B local model is slow (100-200 s per question on my laptop) and sometimes misses a fact one run and finds it the next.
Code
Two files, Python 3.10, pip install mcp, plus Ollama with gemma3:4b. Set SANITY_CONTEXT_MCP_URL and SANITY_API_TOKEN (a Context Viewer token).
kb_client.py - MCP client for the Sanity Context endpoint
"""Tiny client for the Funding Fit Sanity Context MCP endpoint."""
import asyncio, json, os, sys
from contextlib import asynccontextmanager
from pathlib import Path
from mcp import ClientSession
from mcp.client.streamable_http import streamable_http_client
from mcp.shared._httpx_utils import create_mcp_http_client
CFG_PATH = Path('sanity.json') # or set the two env vars
def load_cfg():
"""Env vars first (SANITY_CONTEXT_MCP_URL, SANITY_API_TOKEN), else a local json file."""
url, tok = os.environ.get('SANITY_CONTEXT_MCP_URL'), os.environ.get('SANITY_API_TOKEN')
if url and tok:
return {'mcp_url': url, 'token': tok}
return json.loads(CFG_PATH.read_text(encoding='utf-8-sig'))
@asynccontextmanager
async def session():
cfg = load_cfg()
client = create_mcp_http_client(headers={'Authorization': 'Bearer ' + cfg['token']})
async with streamable_http_client(cfg['mcp_url'], http_client=client) as streams:
async with ClientSession(streams[0], streams[1]) as s:
await s.initialize()
yield s
def text_of(result):
out = []
for c in result.content:
t = getattr(c, 'text', None)
if t:
out.append(t)
return '\n'.join(out)
async def call(s, name, args=None):
r = await s.call_tool(name, args or {})
return text_of(r)
if __name__ == '__main__':
sys.stdout.reconfigure(encoding='utf-8')
async def main():
async with session() as s:
print(await call(s, 'initial_context'))
asyncio.run(main())
funding_fit.py - the agent
"""Funding Fit: which early-stage programs can a solo, non-incorporated founder in Korea apply to?
Agent loop. Every model call runs locally through Ollama; every fact comes from the
Sanity Context MCP endpoint (knowledge base kbbjiFEOCdVH).
1. initial_context -> outline of entry paths
2. local model plans -> which entries to read + search keywords (JSON)
3. knowledge_base_search -> extra recall; knowledge_base_read per entry
4. split entries by '## <program>' sections and group them per program
5. local model extracts per program: field values + a verbatim quote for each (JSON)
6. grounding check: a value counts only if its quote is really in the retrieved text
7. fit verdict is computed in code from the verified fields (no model judgement)
"""
import argparse, asyncio, json, re, sys, time
import urllib.request
from kb_client import session, call
OLLAMA = 'http://127.0.0.1:11434/api/chat'
KB = 'kbbjiFEOCdVH'
MAX_READ = 6
FIELDS = {
'company_required': '법인 필요',
'in_person_required': '현장 참여 필수',
'location': '장소',
'english_video_interview': '영어 화상 면접',
'fee': '참가비',
'check_size': '투자 규모',
'how_to_apply': '지원 방법',
'deadline': '마감',
'application_language': '지원 언어',
'decision_time': '결과까지',
}
YESNO = ('company_required', 'in_person_required', 'english_video_interview', 'fee')
def llm(prompt, model, as_json=False, num_predict=700):
body = {'model': model, 'messages': [{'role': 'user', 'content': prompt}], 'stream': False,
'options': {'temperature': 0, 'repeat_penalty': 1.1, 'num_predict': num_predict, 'num_ctx': 8192}}
if as_json:
body['format'] = 'json'
req = urllib.request.Request(OLLAMA, data=json.dumps(body).encode('utf-8'),
headers={'Content-Type': 'application/json'})
last = None
for _ in range(3):
try:
with urllib.request.urlopen(req, timeout=300) as r:
return json.loads(r.read().decode('utf-8'))['message']['content']
except Exception as e: # local runner hiccup: retry briefly
last = e
time.sleep(2)
raise RuntimeError(f'local model failed: {last}')
def outline_paths(ctx):
tail = ctx.split('entries.', 1)[-1]
return sorted(set(m.group(1) for m in re.finditer(r'(?m)^\s*([a-z0-9][a-z0-9_\-/]*)\s*(?:\[core\])?\s*$', tail)))
def norm(s):
s = s.replace('**', '').replace('’', "'").replace('“', '"').replace('”', '"')
return re.sub(r'\s+', ' ', s).strip().lower()
def program_key(title):
t = re.sub(r'\(.*?\)', '', title).strip().lower()
return re.sub(r'[^a-z0-9가-힣]+', ' ', t).strip()
def split_programs(entry_path, text):
"""Return {program_key: (title, [(entry_path, section_text)])} from '## Title' sections."""
out = {}
parts = re.split(r'(?m)^## ', text)
for p in parts[1:]:
title, _, body = p.partition('\n')
k = program_key(title)
if not k or len(body.strip()) < 40:
continue
out.setdefault(k, (title.strip(), []))[1].append((entry_path, body.strip()))
return out
PLAN = """Pick knowledge base entries to read for this question about startup funding programs.
Valid entry paths: {paths}
Question (may be Korean): {q}
Return JSON {{"paths": [up to {n} paths from the valid list], "keywords": "6-10 English search words"}}"""
EXTRACT = """From the SOURCE about the program "{name}", fill these fields.
For each field give "value" and "quote". The quote MUST be copied word for word from SOURCE (5-25 words).
If SOURCE does not say, use value "unknown" and quote "".
Fields:
- company_required: "yes" if an incorporated company is required to apply, "no" if not required, else "unknown"
- in_person_required: "yes" if in-person/on-site participation or relocation is required, "no" if fully remote/virtual, else "unknown"
- location: city/country of the program, or "unknown"
- english_video_interview: "yes" if interviews or required sessions are video calls in English, "no" if not, else "unknown"
- fee: "yes" if founders must pay tuition or a fee, "no" if stated free, else "unknown"
- check_size: investment amount, or "unknown"
- how_to_apply: form, email or website, or "unknown"
- deadline: date, "rolling", or "unknown"
- application_language: language(s) the application or pitch can be written in (e.g. "Korean", "English"), or "unknown"
- decision_time: how long until a decision or first answer (e.g. "1 business day", "2 weeks"), or "unknown"
Return JSON {{"company_required": {{"value": "..", "quote": ".."}}, ...all 10 fields}}
SOURCE:
{src}"""
# A yes/no value only counts when its quote talks about that topic.
KEYS = {
'company_required': r'incorporat|company|entity|corporation|registered business|법인',
'in_person_required': r'in-person|in person|on-site|onsite|physical|relocat|residen|remote|virtual|full-time',
'english_video_interview': r'video|zoom|virtual|online|call|english',
'fee': r'fee|tuition|cost|free of charge|pay|charge',
'check_size': r'invest|check|up to|per team|per company|safe|cps|krw|₩|usd',
'application_language': r'korean|english|language|한국어|영어',
'decision_time': r'day|week|month|within|decision|review|result|respond|reply',
}
GENERIC = re.compile(r'^(sources?|references?|summary|overview|notes?)$|matrix|comparison|at a glance|calendar|format|requirements|differentiators', re.I)
PROFILE_ENTRIES = ('eligibility', 'korea_programs', 'global_programs')
def verify(field, field_obj, src_norm, src_lines=()):
if not isinstance(field_obj, dict):
return 'unknown', '', False
v = str(field_obj.get('value', 'unknown')).strip() or 'unknown'
q = str(field_obj.get('quote', '')).strip()
if v.lower() == 'unknown':
return 'unknown', '', True
ok = bool(q) and len(q) >= 8 and norm(q) in src_norm
if ok and field in KEYS:
nq = norm(q)
line = next((l for l in src_lines if nq in l), nq) # keep the '- **Incorporation:**' label in view
if not re.search(KEYS[field], line):
ok = False # real quote, but about something else
return v, q, ok
def verdict(rec):
"""Fit for: solo, no company, stays in Korea, no English video interview, no fee."""
blockers, unknown = [], []
rules = [('company_required', 'yes', '법인 필요'), ('in_person_required', 'yes', None),
('english_video_interview', 'yes', '영어 화상 면접'), ('fee', 'yes', '참가비 있음')]
for f, bad, label in rules:
v, _, ok = rec[f]
if not ok:
unknown.append(FIELDS[f])
continue
if v.lower().startswith(bad):
if f == 'in_person_required':
loc = rec['location'][0] if rec['location'][2] else ''
if 'korea' in loc.lower() or 'seoul' in loc.lower():
blockers.append('서울 현장 상주')
continue
label = f'해외 현장 참여({loc or "장소 미상"})'
blockers.append(label)
elif v.lower() == 'unknown':
unknown.append(FIELDS[f])
if any(b.startswith('해외') or b.startswith('법인') or b.startswith('영어') for b in blockers):
return '맞지 않음', blockers, unknown
if blockers:
return '조건부', blockers, unknown
return ('맞음' if not unknown else '가능성 있음(확인 필요)'), blockers, unknown
async def ask(q, model):
trace = []
async with session() as s:
ctx = await call(s, 'initial_context')
trace.append('initial_context')
valid = outline_paths(ctx)
try:
plan = json.loads(llm(PLAN.format(paths=', '.join(valid), q=q, n=MAX_READ), model, True, 200))
except Exception:
plan = {}
paths = [p for p in plan.get('paths', []) if p in valid]
kw = plan.get('keywords') or 'incorporation remote in-person interview language apply'
kw = ' '.join(kw) if isinstance(kw, list) else str(kw)
found = await call(s, 'knowledge_base_search', {'knowledgeBase': KB, 'query': kw, 'return': 'paths'})
trace.append(f'knowledge_base_search({kw!r})')
for p in re.findall(r'`([^`]+)`', found):
if p in valid and p not in paths:
paths.append(p)
for p in reversed(PROFILE_ENTRIES): # program profiles are always read
if p in valid:
if p in paths:
paths.remove(p)
paths.insert(0, p)
paths = paths[:MAX_READ]
entries = {}
for p in paths:
entries[p] = await call(s, 'knowledge_base_read', {'knowledgeBase': KB, 'paths': [p]})
trace.append(f'knowledge_base_read x{len(paths)}')
names = {}
for p in PROFILE_ENTRIES:
for k, (title, _) in split_programs(p, entries.get(p, '')).items():
if not GENERIC.search(title) and k not in names:
names[k] = re.sub(r'\s*\(.*?\)', '', title).strip()
programs = {k: [t, []] for k, t in names.items()}
for p, text in entries.items():
for k, (title, secs) in split_programs(p, text).items():
owner = next((n for n in names if k == n or k.startswith(n + ' ') or n.startswith(k + ' ')), None)
if owner:
programs[owner][1].extend(secs)
programs = {k: v for k, v in programs.items() if v[1]}
results = []
for k, (title, secs) in programs.items():
# one small extraction per section (a 4B model truncates long JSON), then merge:
# a verified value beats an unverified one; profile entries are read first.
rec = {f: ('unknown', '', False) for f in FIELDS}
secs = sorted(secs, key=lambda s: PROFILE_ENTRIES.index(s[0]) if s[0] in PROFILE_ENTRIES else 9)
for ep, body in secs:
src = f'[{ep}]\n{body[:2500]}'
src_norm = norm(src)
src_lines = [norm(l) for l in src.splitlines() if l.strip()]
try:
raw = json.loads(llm(EXTRACT.format(name=title, src=src), model, True, 1400))
except Exception:
raw = {}
for f in FIELDS:
v, qt, ok = verify(f, raw.get(f), src_norm, src_lines)
if v == 'unknown':
continue
if rec[f][0] == 'unknown' or (ok and not rec[f][2]):
rec[f] = (v, qt, ok)
v, blockers, unknown = verdict(rec)
results.append({'program': title, 'entries': sorted(set(ep for ep, _ in secs)), 'fields': rec,
'verdict': v, 'blockers': blockers, 'unknown': unknown})
wants = parse_wants(q)
answer, top = compose(results, wants)
known = {x['program'] for x in results}
return {'trace': trace, 'read': paths, 'results': results, 'answer': answer, 'top': top, 'wants': sorted(wants), 'programs': sorted(known)}
WANT_LABEL = {'domestic': '국내', 'korean': '한국어 지원', 'rolling': '상시 접수', 'fast': '빠른 결과'}
def parse_wants(q):
"""Read extra conditions from the question itself (code, not the model)."""
w = set()
if re.search(r'국내|korea|한국 ?(투자|회사|vc)', q, re.I):
w.add('domestic')
if re.search(r'한국어|in korean', q, re.I):
w.add('korean')
if re.search(r'지금|바로|상시|수시|rolling|year-round', q, re.I):
w.add('rolling')
if re.search(r'빠른|빨리|빠르|fast|quick', q, re.I):
w.add('fast')
return w
UNIT_DAYS = {'hour': 1 / 24, '시간': 1 / 24, 'business day': 1.4, '영업일': 1.4, 'day': 1, '일': 1, 'week': 7, '주': 7, 'month': 30, '개월': 30, '달': 30}
def speed_days(x):
v, _, ok = x['fields']['decision_time']
if not ok or v == 'unknown':
return None
t = v.lower().replace('one ', '1 ').replace('a ', '1 ')
m = re.search(r'(\d+(?:\.\d+)?)[^\d]{0,12}?(hour|시간|business day|영업일|day|week|month|개월|일|주|달)', t)
return float(m.group(1)) * UNIT_DAYS[m.group(2)] if m else None
def matches(x, wants):
f, notes, ok = x['fields'], [], True
loc = f['location'][0].lower() if f['location'][2] and f['location'][0] != 'unknown' else ''
if 'domestic' in wants:
if loc and not re.search(r'korea|seoul|한국|서울', loc):
ok = False
notes.append(f'국내 아님({f["location"][0]})')
elif not loc:
notes.append('위치 근거 없음')
if 'korean' in wants:
lang = f['application_language'][0].lower() if f['application_language'][2] and f['application_language'][0] != 'unknown' else ''
if lang and not re.search(r'korean|한국어', lang):
ok = False
notes.append(f'지원 언어 {f["application_language"][0]}')
elif not lang:
notes.append('지원 언어 근거 없음')
if 'rolling' in wants:
dl = f['deadline'][0].lower() if f['deadline'][2] and f['deadline'][0] != 'unknown' else ''
if dl and not re.search(r'rolling|year-round|anytime|상시|수시', dl):
notes.append(f'마감 {f["deadline"][0]}')
elif not dl:
notes.append('마감 근거 없음')
if 'fast' in wants and speed_days(x) is None:
notes.append('결과 시점 근거 없음')
return ok, notes
def compose(results, wants=frozenset()):
base = compose_base(results)
if not wants:
return base, [x['program'] for x in results if x['verdict'] in ('맞음', '가능성 있음(확인 필요)')]
lines = [base, '', '질문 조건(' + ', '.join(WANT_LABEL[w] for w in sorted(wants)) + ')으로 다시 거른 결과:']
keep, drop = [], []
for x in results:
if x['verdict'] not in ('맞음', '가능성 있음(확인 필요)', '조건부'):
continue
ok, notes = matches(x, wants)
(keep if ok else drop).append((x, notes))
if 'fast' in wants:
keep.sort(key=lambda t: speed_days(t[0]) if speed_days(t[0]) is not None else 1e9)
for x, notes in keep:
f = x['fields']
extra = [f'{FIELDS[k]} {f[k][0]}' for k in ('decision_time', 'application_language', 'deadline') if f[k][2] and f[k][0] != 'unknown']
head = '조건부' if x['verdict'] == '조건부' else '후보'
s = f"- [{head}] {x['program']}"
if extra:
s += f" ({'; '.join(extra)})"
if notes:
s += f" — 확인 필요: {', '.join(notes)}"
lines.append(s)
for x, notes in drop:
lines.append(f"- [조건 밖] {x['program']}: {', '.join(notes)}")
top = [x['program'] for x, _ in keep if x['verdict'] != '조건부']
return '\n'.join(lines), top
def compose_base(results):
"""Final summary written by code from verified facts only (the 4B model invented facts here)."""
out = []
for want, head in (({'맞음', '가능성 있음(확인 필요)'}, '지원 후보'), ({'조건부'}, '조건부'), ({'맞지 않음'}, '제외')):
for x in (r for r in results if r['verdict'] in want):
f = x['fields']
bits = [f'{FIELDS[k]} {f[k][0]}' for k in ('how_to_apply', 'deadline', 'check_size') if f[k][2] and f[k][0] != 'unknown']
why = ', '.join(x['blockers']) if x['blockers'] else '걸리는 조건 없음'
s = f"- [{head}] {x['program']}: {why}"
if bits:
s += f" ({'; '.join(bits)})"
if x['unknown'] and head == '지원 후보':
s += f" — 자료에 없음: {', '.join(x['unknown'])}"
out.append(s)
return '\n'.join(out)
ORDER = {'맞음': 0, '가능성 있음(확인 필요)': 1, '조건부': 2, '맞지 않음': 3}
def render(q, r, secs):
lines = [f'질문: {q}', f'도구 호출: {" -> ".join(r["trace"])}', f'읽은 항목: {", ".join(r["read"])}', '']
res = sorted(r['results'], key=lambda x: ORDER.get(x['verdict'], 9))
checked = verified = 0
for x in res:
lines.append(f'■ {x["program"]} → {x["verdict"]}')
if x['blockers']:
lines.append(f' 걸리는 점: {", ".join(x["blockers"])}')
if x['unknown']:
lines.append(f' 자료에 없거나 근거 불충분: {", ".join(x["unknown"])}')
for f, label in FIELDS.items():
v, quote, ok = x['fields'][f]
if v == 'unknown':
continue
checked += 1
verified += ok
mark = '✓' if ok else '✗ 근거 부족(무시함)'
lines.append(f' - {label}: {v} {mark} “{quote[:90]}”')
lines.append(f' 출처 항목: {", ".join(x["entries"])}')
lines.append('')
lines.append('정리(코드가 확인된 사실만으로 작성):')
lines.append(r.get('answer', ''))
lines.append('')
ok_list = r.get('top') or []
lines.append('한 줄 결론: ' + (', '.join(ok_list) + ' 부터 확인해 보세요.' if ok_list else '조건에 맞는 곳을 자료에서 찾지 못했어요.'))
lines.append(f'근거 확인: 사실 {checked}개 중 {verified}개가 원문 인용과 일치 | {secs:.1f}초')
return '\n'.join(lines)
def main():
sys.stdout.reconfigure(encoding='utf-8')
ap = argparse.ArgumentParser()
ap.add_argument('question', nargs='?', default='법인 없이 혼자 지원할 수 있고, 해외 현장 참여나 영어 화상 면접이 필요 없는 곳은?')
ap.add_argument('--model', default='gemma3:4b')
ap.add_argument('--json', help='also save raw results to this path')
a = ap.parse_args()
t0 = time.time()
r = asyncio.run(ask(a.question, a.model))
print(render(a.question, r, time.time() - t0))
if a.json:
with open(a.json, 'w', encoding='utf-8') as f:
json.dump(r, f, ensure_ascii=False, indent=1)
if __name__ == '__main__':
main()
Thanks for reading. If you run a program that accepts solo, non-incorporated founders from Korea and it's missing here, tell me in the comments and I'll add your official page to the knowledge base.
Top comments (0)