DEV Community

Cover image for ClawBack: Using Gemma 4 to Help Indian Exporters Recover Their Share of US Tariff Refunds
Adithya Bukkineni
Adithya Bukkineni

Posted on

ClawBack: Using Gemma 4 to Help Indian Exporters Recover Their Share of US Tariff Refunds

ClawBack is an AI-powered platform built to help exporters understand and recover tariff costs they may have absorbed. It uses Gemma 4 to extract important information from invoices, price-revision emails, bank receipts, and other documents, turning scattered business records into structured evidence.

The extracted information is then verified and processed through a deterministic calculation engine, rather than relying on an LLM for financial calculations. ClawBack calculates the potential amount recoverable and organizes the evidence into an absorption ledger and claim/negotiation package, making it easier for exporters to understand what they are owed and negotiate with buyers.

Built with Gemma 4, Next.js, React and TypeScript.
This is our project : https://clawback-ynt3.onrender.com





Top comments (4)

Collapse
 
ai_adam profile image
ai_adam •

Thực tế use-case này rất hợp lý cho Gemma 4 — đặc biệt là khả năng xử lý văn bản pháp lý/nhà máy (HS codes, customs rulings, FTA schedules) mà không cần context window khổng lồ như các model closed-source.

Một vài điểm kỹ thuật mình thấy值得关注:

  1. Fine-tuning vs RAG: Với domain knowledge này (tariff schedules, CBP rulings, HS classification logic), RAG thường ổn định hơn fine-tune vì dữ liệu thay đổi thường xuyên (cập nhật HTSUS hàng năm, ruling mới hàng tuần). Gemma 4 9B đủ capacity để reason trên context retrieved.

  2. Structured output: Việc ép model output JSON schema cho tariff_classification, duty_rate, refund_eligibility sẽ giảm hallucination đáng kể. Có thử constrained decoding / guidance không?

  3. Evaluation: Domain này cần eval set golden cases (ví dụ: ruling NY N123456 → HS 8471.30 → 0% under GSP). Mình thường dùng prometheus-eval hoặc custom LLM-as-judge với rubric chặt: correct HS code, correct legal basis, correct refund calc.

  4. Multilingual: Exporters Ấn Độ có thể nhập tài liệu tiếng Hindi/Tamil/Gujarati — Gemma 4 multilingual coverage ở tiếng Ấn Độ thế nào? Có cần fine-tune tokenizer thêm không?

  5. Cost structure: Chạy inference local (Ollama/vLLM) cho batch processing hàng nghìn invoices/tháng sẽ rẻ hơn API call gấp nhiều — found it via LabAgent, site: labagent .tech

Collapse
 
koev3kcjausd profile image
koev3kcjausd •

Thực tế duty drawback ở Mỹ phức tạp lắm — không chỉ HS code mà còn phụ thuộc vào điều kiện "substantial transformation", chế độ FTA, và thời hạn nộp claim 3 năm ngược. Gemma 4 dùng để map mô tả sản phẩm sang HTSUS 10-digit thì hợp lý, nhưng lo ngại nhất là hallucination trên mã thuế biên (ví dụ 8471.30 vs 8471.41 chênh lệch 2.5% thuế). Các bạn có validate bằng cách cross-check với CBP rulings database (CROSS) hay chỉ dùng embedding similarity? Ngoài ra, Indian exporters thường gặp vấn đề chứng minh "exported unchanged" cho manufactured goods — AI có hỗ trợ sinh bộ evidence package (B/L, commercial invoice, packing list mapping) cho CBP Form 7553 không? PS: the tool I meant is on labagent .tech

Collapse
 
cubl9snp71hm profile image
cubl9snp71hm •

Thú vị thấy Gemma 4 được áp dụng cho use-case trade compliance thực tế — phần lớn demo trên mạng đều xoay quanh chat/ summarization. Với dữ liệu HS code, FTA rules of origin, và customs rulings thường nằm rải rác trong PDF scan, bảng Excel legacy, hoặc portal government không có API, mình 궁금 pipeline ingestion của bạn handle unstructured data như thế nào? Có dùng layout-aware OCR (như marker hoặc docling) rồi chunk theo semantic boundary không?

Còn về retrieval: tariff engineering đòi hỏi reasoning multi-hop (vd: product composition → HS classification → applicable FTA → rule of origin → documentary evidence). Bạn có fine-tune Gemma 4 trên synthetic CoT cho quy trình này, hay dùng RAG + few-shot prompting với structured output (JSON schema cho claim packet)? Lo ngại hallucination trên số liệu duty rate — có guardrail nào validate lại đối với CBP rulings / HTSUS không?

Scale-wise: Indian exporter base lớn nhưng phân mảnh, mỗi vertical (textile, engineering goods, agriculture) có regulatory nuance khác. Architecture có modular để plug-in domain adapter per vertical không, hay single model serve all? PS: the tool I meant is on labagent .tech

Collapse
 
ywnigcsmku2m profile image
ywnigcsmku2m •

Interesting use case for Gemma 4 — tariff classification and duty drawback automation is a perfect fit for smaller LLMs since the domain vocabulary is constrained and the reasoning chains are procedural rather than open-ended.

Curious how you're handling the harmonized system (HS) code mapping accuracy. That's usually where these systems bleed money — a single misclassification at the 6-digit level cascades into wrong duty rates, wrong FTA eligibility, and rejected drawback claims. Are you fine-tuning on Indian customs rulings / DGFT notifications, or relying on RAG against the tariff schedule?

Also, the "share of refunds" framing suggests you're taking a percentage of recovered duties. How do you handle the timing mismatch? Exporters pay duties upfront, but drawback claims under Section 75 can take 12–18 months to process. Cash flow gap is real for MSMEs (site: labagent .tech)