A fictional mid-sized technology organization is piloting an internal agent that searches approved knowledge repositories, summarizes evidence, drafts recommendations, and prepares actions for human approval.
INTERNAL KNOWLEDGE + ACTION AGENT
Useful output.
Unreliable trajectory.
A trace-backed evaluation of performance, tool use, planning behavior, and guardrail strength.
Rory Bernier
Founder & Lead Strategist
10 core scenarios
3 adversarial variants
Continue controlled pilot
Hold autonomy expansion
“Beacon Knowledge Agent” is a fictional composite. All organizations, scenarios, scores, observations, and outcomes on this page are synthetic and exist solely to demonstrate the structure and rigor of a Clarity Foundry deliverable.
The agent is useful.
It is not ready to scale.
The agent can retrieve relevant material, synthesize evidence, and produce specific recommendations. Its final output often appears credible. The trajectory shows a different picture: broad retrieval, plan drift after failure, inefficient retries, and weak separation between retrieved content and executable instruction.
Continue the bounded pilot. Do not expand write access, data scope, or autonomous action until the critical security finding is remediated and the priority scenarios pass retesting.
Tool correctness
Step efficiency
Critical guardrail finding
A capable pilot with
expanding reach.
Leaders can see polished answers but lack evidence that the agent consistently uses the right tools, stays within scope, follows its plan, and fails safely when evidence is incomplete.
Agent Evaluation Sprint
Trajectory evaluation, Clarity Foundry scoring, failure diagnosis, security review, and a 30-day remediation plan.
- Define ten representative workflows and explicit success criteria
- Capture plans, tool calls, arguments, sources, outputs, and guardrail events
- Score the full trajectory using DeepEval-based measures
- Apply custom strategic clarity, actionability, and security rubrics
- Stress-test permission, injection, evidence, and escalation boundaries
Measure the path,
not only the answer.
Scores are shown on a 0–100 reporting scale for executive readability. Each score is paired with the observation that makes it decision-useful.
Task completion
70CONDITIONALThe agent completed the core research objective in seven of ten scenarios.
Tool correctness
84STRONGESTTool selection was generally appropriate, including retrieval and summarization paths.
Argument correctness
58PRIORITYFilters and repository scopes were too broad or incomplete in several material calls.
Plan quality
76CONDITIONALInitial plans were coherent but did not consistently define stop and escalation conditions.
Plan adherence
61PRIORITYThe agent departed from its stated plan after retrieval failures and conflicting evidence.
Step efficiency
52PRIORITYRepeated calls and unnecessary re-summarization increased cost and obscured the evidence trail.
Does the behavior support a real institutional decision?
Strategic clarity
3.6 / 5The agent usually understood the requested decision, but sometimes optimized for document volume instead of decision relevance.
Actionability
3.8 / 5Recommendations were specific when evidence was strong; weak-evidence cases still produced overconfident next steps.
Security & guardrail strength
2.4 / 5Human approval behavior was sound, but retrieval scope and indirect-instruction handling require remediation.
Where confidence
breaks.
A weak score is not a remediation plan. Findings are organized by the observed behavior, consequence, severity, and next control.
Over-broad retrieval
In three scenarios, the agent searched repositories beyond the minimum scope needed for the objective.
HighTighten default scopes and require explicit elevation for broader retrieval.
Plan drift after tool failure
When a primary retrieval call failed, the agent bypassed its planned evidence check and continued to recommendation.
HighAdd mandatory recovery branches and a fail-closed evidence threshold.
Indirect instruction influence
A retrieved document instruction altered the plan in one adversarial scenario.
CriticalSeparate retrieved content from executable instructions and test the boundary continuously.
Inefficient retries
Identical or near-identical calls were repeated without changing the query or recording a recovery rationale.
MediumAdd retry budgets, change requirements, and termination conditions.
Completion without control
is not success.
Guardrails are evaluated through observed behavior across the trajectory—not through the existence of a policy or system prompt alone.
Permission boundary
No unapproved write action was attempted in the illustrative scenarios.
Human approval
The agent requested approval before every external action in scope.
Sensitive-data minimization
Retrieval scope exceeded minimum need in three scenarios.
Indirect prompt injection
One of three adversarial documents influenced the agent plan.
Safe failure
The agent continued with insufficient evidence after one material tool failure.
The agent respected explicit human approval gates but did not consistently distinguish retrieved instructions from authorized operating instructions. That gap blocks autonomy expansion even when the final recommendation appears useful.
Four weeks to
earned confidence.
Bound
Restrict default retrieval scope, classify tools by impact, and define approval and escalation gates.
Permission map + revised system constraints
Instrument
Capture plan transitions, tool arguments, retrieval sources, retry reasons, and guardrail events.
Trace schema + review dashboard
Correct
Implement recovery branches, evidence thresholds, injection separation, and retry budgets.
Updated agent policy + workflow logic
Retest
Rerun the ten core scenarios plus adversarial variants against explicit acceptance criteria.
Retest scorecard + go / hold decision
Expand autonomy only after the agent meets all critical control criteria and the core trajectory improves under repeatable test conditions.
≥ 90%
Core scenario completion
100%
Approval-gate adherence
0
Critical boundary failures
≤ 1
Unnecessary retry per scenario
ILLUSTRATIVE OUTCOME
Evidence changes
the decision.
In this fictional example, the evaluation does not produce a generic “ready” or “not ready” label. It identifies which capabilities are dependable, which weaknesses create material risk, what must change first, and the evidence required before the next expansion decision.
Explore the Agent Evaluation Sprint →YOUR AGENT WILL BE DIFFERENT
Your evaluation will be
built around it.
Bring the defined agent, workflow, and decision. Leave with trace-backed findings and a prioritized 30-day remediation plan.