Description
Revalvo is a local-first workbench designed for evaluating large language models (LLMs) by running prompts across multiple models simultaneously. It provides a comprehensive environment where users can write, run, version, and evaluate prompts before they reach production. Unlike traditional chat UIs, Revalvo offers full control over prompts, models, and variables, ensuring transparency and traceability with scores, diffs, and version history.
The platform is not a typical SaaS; it requires no account creation, and all data remains in the user's browser, ensuring privacy and security. Users can connect their providers using keys, OAuth, or local runtimes, and create separate workspaces for each provider. The playground feature allows users to write once and run everywhere, with models responding to the same prompt in parallel, providing metrics such as latency, tokens, and cost per response.
Revalvo's version control system treats prompts like code, with immutable snapshots, Git-style version history, and the ability to roll back to any previous version. Users can export versions as API snippets or agent prompts and duplicate prompts across workspaces for fresh evaluations.
The platform supports large-scale testing with datasets, offering 40 batch-runnable scorers for structure, AI judgment, performance, and safety. Reports generated from batch runs include pass rates, latency, and cost per model, providing a detailed receipt of readiness for shipping.
Revalvo is optimized for efficiency, with thoughtful defaults and minimal setup time. It balances the speed of chat playgrounds with the rigor of full evaluation platforms, making it suitable for both first-time users and experienced prompt engineers.
Revalvo's Core Features
Local-first LLM evaluation
Real-time scoring and version control
No account required
Full transparency with scores and diffs
Connect via keys, OAuth, or local runtimes
Parallel model response evaluation
Git-style version history
Batch testing with 40 scorers
How to use Revalvo?
Configure: Connect your provider using keys or OAuth
Use: Run prompts across multiple models in parallel
Optimise: Save and version your prompts with Git-style history
Test: Run batch tests with datasets and scorers
Revalvo's Use Cases
- Prompt Evaluation
- Version Control
- Batch Testing
- Model Comparison
- Data Privacy





