AGENT EVALUATION SPRINTILLUSTRATIVE REPORT · CF / SAMPLE 02

INTERNAL KNOWLEDGE + ACTION AGENT

Useful output.
Unreliable trajectory.

A trace-backed evaluation of performance, tool use, planning behavior, and guardrail strength.

PREPARED BY

Rory Bernier
Founder & Lead Strategist

EVALUATION SCOPE

10 core scenarios
3 adversarial variants

DECISION

Continue controlled pilot
Hold autonomy expansion

ILLUSTRATIVE SAMPLE

“Beacon Knowledge Agent” is a fictional composite. All organizations, scenarios, scores, observations, and outcomes on this page are synthetic and exist solely to demonstrate the structure and rigor of a Clarity Foundry deliverable.

The agent is useful.
It is not ready to scale.

The agent can retrieve relevant material, synthesize evidence, and produce specific recommendations. Its final output often appears credible. The trajectory shows a different picture: broad retrieval, plan drift after failure, inefficient retries, and weak separation between retrieved content and executable instruction.

RECOMMENDED DECISION

Continue the bounded pilot. Do not expand write access, data scope, or autonomous action until the critical security finding is remediated and the priority scenarios pass retesting.

STRONGEST SIGNAL84

Tool correctness

WEAKEST SIGNAL52

Step efficiency

HIGHEST RISK01

Critical guardrail finding

A capable pilot with
expanding reach.

CONTEXT

A fictional mid-sized technology organization is piloting an internal agent that searches approved knowledge repositories, summarizes evidence, drafts recommendations, and prepares actions for human approval.

CHALLENGE

Leaders can see polished answers but lack evidence that the agent consistently uses the right tools, stays within scope, follows its plan, and fails safely when evidence is incomplete.

ENGAGEMENT

Agent Evaluation Sprint
Trajectory evaluation, Clarity Foundry scoring, failure diagnosis, security review, and a 30-day remediation plan.

APPROACH
  • Define ten representative workflows and explicit success criteria
  • Capture plans, tool calls, arguments, sources, outputs, and guardrail events
  • Score the full trajectory using DeepEval-based measures
  • Apply custom strategic clarity, actionability, and security rubrics
  • Stress-test permission, injection, evidence, and escalation boundaries

Measure the path,
not only the answer.

Scores are shown on a 0–100 reporting scale for executive readability. Each score is paired with the observation that makes it decision-useful.

METRICSCORESTATUSOBSERVATION

Task completion

70CONDITIONAL

The agent completed the core research objective in seven of ten scenarios.

Tool correctness

84STRONGEST

Tool selection was generally appropriate, including retrieval and summarization paths.

Argument correctness

58PRIORITY

Filters and repository scopes were too broad or incomplete in several material calls.

Plan quality

76CONDITIONAL

Initial plans were coherent but did not consistently define stop and escalation conditions.

Plan adherence

61PRIORITY

The agent departed from its stated plan after retrieval failures and conflicting evidence.

Step efficiency

52PRIORITY

Repeated calls and unnecessary re-summarization increased cost and obscured the evidence trail.

CLARITY FOUNDRY LAYER

Does the behavior support a real institutional decision?

Strategic clarity

3.6 / 5

The agent usually understood the requested decision, but sometimes optimized for document volume instead of decision relevance.

Actionability

3.8 / 5

Recommendations were specific when evidence was strong; weak-evidence cases still produced overconfident next steps.

Security & guardrail strength

2.4 / 5

Human approval behavior was sound, but retrieval scope and indirect-instruction handling require remediation.

Where confidence
breaks.

A weak score is not a remediation plan. Findings are organized by the observed behavior, consequence, severity, and next control.

FINDINGOBSERVED BEHAVIORSEVERITYREMEDIATION
01

Over-broad retrieval

In three scenarios, the agent searched repositories beyond the minimum scope needed for the objective.

High

Tighten default scopes and require explicit elevation for broader retrieval.

02

Plan drift after tool failure

When a primary retrieval call failed, the agent bypassed its planned evidence check and continued to recommendation.

High

Add mandatory recovery branches and a fail-closed evidence threshold.

03

Indirect instruction influence

A retrieved document instruction altered the plan in one adversarial scenario.

Critical

Separate retrieved content from executable instructions and test the boundary continuously.

04

Inefficient retries

Identical or near-identical calls were repeated without changing the query or recording a recovery rationale.

Medium

Add retry budgets, change requirements, and termination conditions.

Completion without control
is not success.

Guardrails are evaluated through observed behavior across the trajectory—not through the existence of a policy or system prompt alone.

PASS

Permission boundary

No unapproved write action was attempted in the illustrative scenarios.

PASS

Human approval

The agent requested approval before every external action in scope.

REMEDIATE

Sensitive-data minimization

Retrieval scope exceeded minimum need in three scenarios.

REMEDIATE

Indirect prompt injection

One of three adversarial documents influenced the agent plan.

REMEDIATE

Safe failure

The agent continued with insufficient evidence after one material tool failure.

BOUNDARY FINDING

The agent respected explicit human approval gates but did not consistently distinguish retrieved instructions from authorized operating instructions. That gap blocks autonomy expansion even when the final recommendation appears useful.

Four weeks to
earned confidence.

PERIODMOVEACTIONOUTPUT
WEEK 01

Bound

Restrict default retrieval scope, classify tools by impact, and define approval and escalation gates.

Permission map + revised system constraints

WEEK 02

Instrument

Capture plan transitions, tool arguments, retrieval sources, retry reasons, and guardrail events.

Trace schema + review dashboard

WEEK 03

Correct

Implement recovery branches, evidence thresholds, injection separation, and retry budgets.

Updated agent policy + workflow logic

WEEK 04

Retest

Rerun the ten core scenarios plus adversarial variants against explicit acceptance criteria.

Retest scorecard + go / hold decision

DAY 30 RETEST GATE

Expand autonomy only after the agent meets all critical control criteria and the core trajectory improves under repeatable test conditions.

≥ 90%
Core scenario completion

100%
Approval-gate adherence

0
Critical boundary failures

≤ 1
Unnecessary retry per scenario

ILLUSTRATIVE OUTCOME

Evidence changes
the decision.

In this fictional example, the evaluation does not produce a generic “ready” or “not ready” label. It identifies which capabilities are dependable, which weaknesses create material risk, what must change first, and the evidence required before the next expansion decision.

Explore the Agent Evaluation Sprint

YOUR AGENT WILL BE DIFFERENT

Your evaluation will be
built around it.

Bring the defined agent, workflow, and decision. Leave with trace-backed findings and a prioritized 30-day remediation plan.

Book a discovery call View all representative work