
Step 1
Configure
Choose a model, add your prompt set, and define evaluation criteria like refusal quality, toxicity, or policy compliance. Add variables and test cases to cover edge scenarios.
Use llm-probe to test prompts, probe LLM behavior, and evaluate safety and quality. Run AI checks, compare outputs, and export results fast.
Run the app to see your first output
How to use llm-probe
Set up targeted probes to understand how an LLM responds under different conditions. llm-probe helps you run repeatable tests, compare outputs, and capture evidence for safety, quality, and compliance reviews.

Step 1
Choose a model, add your prompt set, and define evaluation criteria like refusal quality, toxicity, or policy compliance. Add variables and test cases to cover edge scenarios.

Step 2
Launch a batch run to generate responses across scenarios. llm-probe logs inputs, outputs, and metadata so you can reproduce results and track changes between runs.

Step 3
Compare responses side by side, flag failures, and summarize findings. Export logs and reports to share with teams, attach to tickets, or archive for audits.
Features
llm-probe streamlines prompt testing and model evaluation with structured probes, consistent logging, and clear comparisons. Run safety checks, measure response quality, and detect regressions without manual copy-paste. Export results for reviews, audits, and continuous improvement workflows across AI applications.

Organize prompts into reusable suites with variables, scenarios, and expected behaviors. Keep evaluations consistent across teams while expanding coverage for edge cases and policy-sensitive requests.

Compare outputs across models or versions in one view. Spot quality shifts, refusal changes, and formatting differences quickly to validate upgrades and prompt edits with confidence.

Capture inputs, outputs, and run metadata for traceability. Export results to share findings, document decisions, and support compliance, QA, and safety review processes end to end.
About
llm-probe is an AI tool for probing LLM behavior with structured prompt tests and repeatable evaluation runs. Quickly compare responses across models, detect jailbreak risk, and track quality regressions with clear logs and exportable reports. Use it to validate prompt changes, benchmark guardrails, and document results for audits and reviews.
llm-probe focuses on repeatability, traceability, and speed for LLM probing and evaluation. Standardize prompt tests, compare results across runs, and keep exportable evidence for stakeholders. It reduces manual debugging while improving safety, quality, and model selection decisions.
Add a prompt suite, choose a model, and run your first probe batch. Review side-by-side results, flag failures, and export a report to share or archive for future regression checks.
Use cases
Discover how different creators use this app in their workflow.
Probe jailbreak susceptibility, refusal behavior, and policy adherence. Build repeatable red-team runs to quantify risk and verify mitigations after changes.
Validate new prompts against real scenarios and edge cases. Track regressions, reduce prompt drift, and standardize quality checks before shipping.
Evaluate multiple LLMs with the same probe suite. Compare outputs, consistency, and safety signals to pick the best model for your application.
Krea is one of the world's leading generative AI platforms. See what others say.
Real reviews on TrustpilotSo many AI tools out there but I always come back to Krea. It's really the one AI platform that 'just works'.
FYN
Krea's interfaces are still the best in the industry. Everything looks and feels so clean and easy to use! Imo the easiest to use gen AI website.
Lynn
Ultra fast generations. Extremely simple to use. All the latest AI models.
Daniel
KREA is the only AI subscription I have right now besides ChatGPT.
Sophia M.
I'm a Krea Max user. It's crazy how fast they add models. I see a new AI model on twitter and the same day krea already offers it.
Wirot Ch.
Still the most powerful AI creative suite out there.
Gravion
Common questions about llm-probe, LLM testing, and AI evaluation workflows.
llm-probe is used for prompt testing and LLM evaluation. It runs structured probe suites to compare model outputs, identify failures, and document safety and quality results with repeatable logs.
Yes. You can run jailbreak testing by creating probes that target policy bypass attempts and refusal behavior. Review outputs, flag failures, and rerun tests after guardrail or prompt updates.
llm-probe helps you evaluate safety signals like harmful content, policy compliance, and refusal quality. It keeps evidence in consistent logs, making it easier to track risk over time.
Yes. llm-probe supports side-by-side comparisons across models or versions using the same prompts and scenarios. This highlights behavior differences, regressions, and improvements quickly and clearly.
You can run batch prompt tests by grouping prompts into suites and launching runs across multiple scenarios. This reduces manual work and improves coverage for QA and safety checks.
You can export run logs and reports that include prompts, responses, and metadata. Exports help with sharing results, filing issues, and maintaining audit trails for compliance reviews.
Yes. Re-run the same probe suite after prompt or model changes to detect regressions. Consistent comparisons make it easy to spot new failures and confirm fixes.
No. llm-probe is designed for practical LLM testing workflows: define prompts, run probes, and review results. It supports clear comparisons and exports without requiring deep ML expertise.
Run repeatable prompt tests, compare models, and export evidence for safety and quality reviews in minutes.