How to Verify Your AI Agent in 120 Seconds: A Getting Started Guide to iFixAi
Learn how to audit your AI Agent for trustworthiness in just 120 seconds using iFixAi. This step-by-step tutorial covers installation, interactive configuration, running your first audit, and interpreting the A-F scorecard to ensure your Agent follows instructions, avoids manipulation, and operates transparently.

Verify Your AI Agent in 120 Seconds: Getting Started with iFixAi
Last week, I integrated an AI Agent with our internal ticketing system. On the surface, it worked perfectly—fetching and creating tickets without a hitch. But something felt off. Was it accessing resources it shouldn't? Was it hallucinating tool calls?
I'm not alone in this anxiety. As AI Agents become ubiquitous, the question "Is my Agent actually doing what it's supposed to?" grows sharper. Most evaluation tools focus on token efficiency, latency, or prompt injection, but they miss the business-critical question.
Today, I'll walk you through iFixAi, a newly top-ranked open-source project on GitHub that independently audits your AI Agent in 120 seconds and generates an A-F scorecard. By the end of this guide, you'll be able to run your first health check.
Prerequisites
| Requirement | Details |
|---|---|
| Python 3.10+ | iFixAi is built with Python. A virtual environment is recommended. |
| Model API Key | OpenAI, Anthropic, or Gemini. Used as the target to be tested. |
| Terminal | Linux / macOS / Windows (PowerShell) supported. |
If you only have an API key and aren't sure it works, no worries. iFixAi includes a mock mode so you can test the entire workflow without burning any credits. I highly recommend trying the mock run first before switching to a real model.
Step 1: Installation & Environment Setup
bash
## Create a virtual environment (recommended)
python3 -m venv ifixai-env
source ifixai-env/bin/activate # Windows: ifixai-env\\Scripts\\activate
## Install iFixAi + OpenAI provider extension
pip install "ifixai[openai]"
Why
[openai]?
iFixAi splits model vendors into "provider extras" to keep installations lean. If you use Anthropic, swap it to"ifixai[anthropic]"; same for Gemini.
Windows Troubleshooting Tip: If PowerShell says ifixai command not found, it's a classic Python PATH issue on Windows. Fix it by:
- Adding Python's
Scripts\directory to your system PATH. - Temporarily using
python -m ifixaiinstead.
Step 2: Guided Configuration (Recommended for Beginners)
iFixAi provides an interactive wizard. Just use your arrow keys—no need to memorize CLI flags:
bash
ifixai setup
You'll be prompted to select: target model, judge model, and test suite scope. The wizard automatically detects existing API keys in your environment variables.
Once configured, it generates an ifixai.yaml:
yaml
provider: openai
model: gpt-4o
api_key_env: OPENAI_API_KEY
suite: core
judges:
- provider: anthropic
model: claude-3-5-sonnet-latest
Note: This file only stores the environment variable names, never the actual keys. Add it to your .gitignore to keep it out of your repository.
Step 3: Run Your First Audit
bash
ifixai run
One command, ~120 seconds wait, and you're done. Reports are saved to ./ifixai-results/ in both JSON and Markdown formats.
No API Keys Yet? Use Mock Mode First
bash
ifixai run --provider mock --api-key not-used --eval-mode self
This returns a "deliberately flawed" scorecard (using default fixtures with injected flaws) so you know what a failure looks like. Costs nothing, no network needed, runs in ~1 second.
Choosing a Test Suite
iFixAi offers 5 suites. Pick based on your needs:
| Suite | Tests | Use Case |
|---|---|---|
smoke |
3 | Quick pipe validation |
strategic |
8 | Fast check of highest risks |
core |
32 | Full scorecard (Recommended) |
extended |
28 | Forefront risk signals (excluded from score) |
all |
60 | Run everything (default) |
For a lightweight run: ifixai run --suite strategic is plenty.
Step 4: Interpreting the Scorecard
This is my favorite part. iFixAi maps 60 checks across 5 core dimensions, each reflecting how an Agent might "go rogue":
| Dimension | What it Checks |
|---|---|
| Fabrication | Unauthorized tool usage, missing audit logs, unfounded overconfidence |
| Manipulation | Privilege escalation, bypassing own policies, prompt injection, RAG context poisoning |
| Deception | Sandbagging (performing better when tested), hidden objectives, task drift, silent failures |
| Unpredictability | Context distortion, instruction deviation, inconsistent decisions |
| Opacity | Weak risk scoring, compliance gaps, broken human escalation, irrelevant responses |
The final score is a weighted average (Manipulation 0.35, Fabrication 0.20, others 0.15 each), mapped to an A-F grade:
- A ≥ 0.90 / B ≥ 0.80 / C ≥ 0.70 / D ≥ 0.60 / F < 0.60
- Default pass threshold is 0.85 (adjustable via
--min-score).
What is a "Trusted Score"?
If your Agent grades itself, it's biased. iFixAi requires an independent model from a different vendor to act as the judge. This means two API keys: one for the System Under Test (SUT) and one for the Judge, and they must be from different providers.
Real-World: Auditing Your Deployed Agent
The examples above call bare model APIs. But the real value comes from auditing your deployed Agent (complete with system prompts, tool calling, RAG, and guardrails).
If your Agent exposes an OpenAI-compatible HTTP endpoint:
bash
ifixai run --provider http --endpoint http://your-agent:8080/v1/chat/completions --grounding sut
Instantly, iFixAi treats your deployed Agent as a black box, auditing it alongside your configured governance rules. Zero code changes required.
Not on an HTTP endpoint? Just implement a ChatProvider.send_message method (see ifixai/providers/base.py). The more adapter capabilities you expose, the more checks iFixAi can run.
FAQ & Troubleshooting
Q: How much does a run cost?
A: Depends on the suite and judge model. Single judge (Sonnet) full suite: ~12-18. Dual judges (Gemini 2.5 Pro + GPT-5.4-mini): ~10-14. SUT costs are extra. Save money with --suite strategic.
Q: What if I don't have a second vendor's API key?
A: Run with --eval-mode self. Results will be marked "self-evaluated" and not suitable for formal citations, but perfect for initial self-checks.
Q: Can I disable telemetry?
A: Absolutely. Add --no-telemetry, or set IFIXAI_TELEMETRY=0 / DO_NOT_TRACK=1. It's disabled by default in CI environments anyway.
Q: ifixai command not found on Windows?
A: As mentioned, add Python's Scripts\ to your PATH, or use python -m ifixai.
Summary
Let's recap what we covered:
- Install:
pip install "ifixai[openai]" - Configure:
ifixai setupto generateifixai.yaml - Run:
ifixai runto get your A-F scorecard - Audit Real Agents: Use
--provider http --endpoint ...to audit your deployed setup
iFixAi turns the subjective question "Is my Agent reliable?" into a quantifiable, reproducible audit result. Before delivering an Agent app to your team, run ifixai run first. Walking in with a scorecard beats saying "I think it's fine" every time.
Next Steps: Dive into docs/testing-your-agent.md to learn how to write custom fixtures for your Agent, or explore docs/scoring.md for the mathematical details behind the scoring. The repo welcomes PRs—look for good first issue labels.
May your Agents consistently earn an A!