A Streamlit tool to evaluate and compare LLM responses across 5 RLHF-inspired quality dimensions — Instruction Following, Truthfulness, Prompt Correctness, Writing Quality, and Verbosity.
-
Updated
Mar 21, 2026 - Python
A Streamlit tool to evaluate and compare LLM responses across 5 RLHF-inspired quality dimensions — Instruction Following, Truthfulness, Prompt Correctness, Writing Quality, and Verbosity.
To associate your repository with the human-feedb topic, visit your repo's landing page and select "manage topics."