Show HN: JevBench, a reproducible benchmark for typed decision models (benchmarkheaven.com)

37 points by florianstandhar 9 hours ago

4 comments:

by sean_pedersen 27 minutes ago

Good project but this one also exists https://huggingface.co/spaces/multimodalart/jev-decision-ind... and the results do not seem to add up and also model sets are different... still needs time to mature likely

by swyx an hour ago
by jldugger 26 minutes ago

Interesting; was curious how this didn't fall into trouble with ToS. Apparently the "no benchmarks" clause was intended for "limited preview" audiences and didn't get removed at launch on accident.

by nzoschke 20 minutes ago

https://is-it-ai-slop.app.mintapis.com/ is a fun tool. Is the source or methodology for that in the github repo? I couldn't find it immediately.

We've been experimenting with Jev for classifying email, some thoughts here: https://housecat.com/blog/classifying-email

Flagging AI written email is a much requested feature too.

Data from: Hacker News, provided by Hacker News (unofficial) API