DEV Community

Cover image for WaterSheep: an open-source alternative to Jev
Samrat Dutta
Samrat Dutta

Posted on AI-assisted

WaterSheep: an open-source alternative to Jev

TypeSafe's Jev made a simple idea popular: a model that doesn't write text, only makes decisions. You send some state and typed questions, and you get back choices, scores and probabilities your code can act on. Jev itself is closed, and you pay per token.

When it launched, I saw everyone hyping it up and thought: I can build an open one. So I did. WaterSheep is an open-source alternative to Jev. It answers the same kinds of questions about any text:

  • noul: yes/no, with the probability of yes
  • choice: one of your options
  • score: a rating scale, with the expected level
  • multi: every option that applies (Jev doesn't have this one)

Every answer includes a probability for every option.

Already using Jev? Keep your code

Run WaterSheep as a local server:

pip install git+https://github.com/SamratDuttaOfficial/WaterSheep
watersheep --model samratduttaofficial/WaterSheep --serve
Enter fullscreen mode Exit fullscreen mode

It answers POST /v1/systemone with the same request and response shape as Jev, so TypeSafe's own Python SDK works against it. Point the SDK at it and run your code as usual:

export TYPESAFE_BASE_URL=http://127.0.0.1:8766
python triage.py
Enter fullscreen mode Exit fullscreen mode

Jev's Python SDK running unchanged against WaterSheep

The script is plain Jev code:

from typesafe_sdk import Choice, Noul, Score, TypeSafeClient

client = TypeSafeClient()  # or TypeSafeClient(base_url="http://127.0.0.1:8766")

r = client.system_one(
    state={"ticket": "I was charged twice for order #4411. Refund the duplicate today."},
    questions={
        "refund": Noul(instructions="Does the customer ask for a refund?"),
        "team": Choice(
            instructions="Which team should handle this?",
            criteria={"billing": None, "shipping": None, "technical": None},
        ),
        "urgency": Score(
            instructions="How urgent is this?",
            criteria=["low", "medium", "high"],
        ),
    },
)
a = r.answers
print("refund:", a["refund"].noul)                               # 0.9451
print("team:", a["team"].choice, a["team"].probabilities)        # billing {'billing': 0.8817, ...}
print("urgency:", a["urgency"].score, a["urgency"].legend)       # 1.4805 {0: 'low', 1: 'medium', 2: 'high'}
Enter fullscreen mode Exit fullscreen mode

Any API key value works locally. I tested this with typesafe-sdk 0.7.2; Jev's JavaScript SDK is untested so far.

No SDK? It's plain HTTP

curl http://127.0.0.1:8766/v1/systemone -d '{
  "state": "My card was charged twice and the box arrived crushed.",
  "questions": {
    "team": {"type": "choice", "instructions": "Which team should handle this?",
             "criteria": {"billing": "payments, refunds", "shipping": "delivery, damage"}},
    "issues": {"type": "multi", "instructions": "Which problems are reported?",
               "criteria": ["double charge", "damaged item", "late delivery"]}
  }
}'
Enter fullscreen mode Exit fullscreen mode

The multi question returns both problems, each with its own probability. A single choice has to pick one.

Python without a server

from transformers import pipeline

ws = pipeline(model="samratduttaofficial/WaterSheep", trust_remote_code=True)
ws("I was charged twice.", question="Which team should handle this?", options=["billing", "shipping", "support"])
Enter fullscreen mode Exit fullscreen mode

It also runs as ONNX, and the live demo lets you try it without installing anything.

How good is it?

On the in-distribution test split it gets 77.8% accuracy; on held-out datasets it never saw, 61.2%. The probabilities are calibrated (expected calibration error 0.026 and 0.043), so you can set thresholds on them: act automatically when the model is confident, and send the rest to a person. Per-benchmark results, including the weak ones, are in the README, and the details are in the paper.

Limits: English only, long inputs are truncated, rating questions are the weakest type, and it isn't meant for high-stakes decisions on its own.

Open all the way down

Code, weights and the training pipeline are Apache 2.0, and you can retrain it on your own data with one script. WaterSheep is independent and not affiliated with TypeSafe.

I'd love feedback, especially examples where it gets things wrong.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to