Candidate application · AI Knowledge / Payroll Agent

Train the system.
Earn the trust.

A working point of view on how a Payroll Agent should learn: grounded in product truth, evaluated against real questions, and improved by every miss.

PAYROLL_AGENT / LEARNING_CORE TRAINING READY
01Sources
02Policies
03Evals
04Feedback
Knowledgeguessing
APPLICATION / 01FLORIDA · 2026

The thesis

A Payroll Agent is only as good as the system that teaches it what is true.

The work is not “write more FAQs.” It is product coverage, knowledge architecture, evaluation, release management, and continuous learning—operated as one system.

Operating principleLanguage can be probabilistic.
Payroll decisions cannot.

The learning curve

How the agent earns maturity.

Interactive model / select a stage
TRAINING MATURITY18%

STAGE 01 / 05

Map reality before teaching answers.

Inventory product flows, payroll concepts, recurring questions, policy owners, and the exact sources the agent is allowed to trust.

Input
Product surface + real questions
Output
Coverage map + gap register
Release gate
Every answerable intent has an owner

The operating loop

From a bad answer to a better system.

Each failure becomes structured evidence. Each fix has a measurable effect. Nothing improves by folklore.

01

Observe

Collect unanswered questions, weak citations, escalations, negative feedback, and product changes.

02

Diagnose

Separate a missing fact from a retrieval miss, reasoning error, stale policy, or unclear product behavior.

03

Teach

Create the smallest durable artifact: a canonical page, worked example, decision tree, or revised metadata.

04

Test

Run gold, edge, ambiguity, jurisdiction, and regression cases. Verify citations and safe abstention.

05

Release

Version the change, record approval, roll out deliberately, and preserve the previous known-good state.

06

Measure

Track coverage, resolution quality, escalation precision, freshness, and the failures that recur.

Non-negotiables

Trust is a product feature.

01

Source-bound

Answers cite approved, current, tenant-aware material.

02

Approval-gated

Anything that can alter pay stops for explicit human review.

03

Privacy-shaped

PII stays inside the narrowest possible authorized boundary.

04

Auditable

Source, reasoning path, action, reviewer, and version are known.

What I would own

The knowledge operation behind the agent.

Clear artifacts turn individual judgment into a durable team capability.

01

Coverage map

Every product capability, intent, source, owner, and known gap.

02

Gold set

Realistic questions with expected facts, citations, actions, and escalation.

03

Failure taxonomy

A shared language for why the agent missed and what kind of fix it needs.

04

Release ledger

What changed, why, who approved it, how it scored, and how to roll it back.

05

Freshness system

Effective dates, locale, jurisdiction, review cadence, and stale-source alerts.

06

Feedback loop

Production evidence converted into a prioritized, measurable learning queue.

Evidence, not adjectives

The work already has
the right shape.

Claude's read-only audit describes an operator who writes long specifications, delegates deeply, uses subagents, and ships from inside the loop. A local workforce-management archive adds direct experience with staffing, time, titles, salaries, budgets, and reporting.

Open the primary audit
CLAUDE / FORENSIC PASS98-DAY WINDOW

“You are dispatching work, not chatting.”

28/28recent active days
66,143tool calls observed
15.0tool calls per turn
3.27%tool error rate
Figures reproduced from the primary Claude report. The source explicitly distinguishes measured local data from inference.
LOCAL ARCHIVE / WORKFORCE MODELLEGACY SYSTEM
Legacy monthly workforce cost report interface
Before agents: custom workforce and payroll reporting. The interface is old; the domain model is not.
VENTURE IN ACTIVE DEVELOPMENTMakeLane.ai

The builder behind the application

From one sentence
to open for business.

MakeLane.ai is the studio Arthur is building to turn an idea into a working revenue system—brand, leads, bookings, deposits, CRM, and follow-up in one connected operating lane. It is living proof of the same discipline this application argues for: frame the system, train the tools, inspect the output, and ship.

Enter MakeLane.ai Studio and builds are advancing in public.
01 / PRODUCT

One connected promise.

Translate a short business idea into a coherent customer journey.

02 / SYSTEMS

Agents with a job to do.

Orchestrate specialized tools around explicit outcomes and gates.

03 / DELIVERY

Working software, not theater.

Move from positioning to production with visible, testable progress.

Codex / professional assessment

What the evidence says about the operator.

Independent synthesis of the available evidence
Arthur reads as a high-agency systems operator: unusually comfortable moving between product thinking, technical execution, and the unglamorous work of making a process run end to end.
01 / Strongest signal

He closes the loop.

The strongest evidence is not raw AI volume. It is the repeated build → verify → ship pattern, plus a low interrupt rate. He gives systems enough context to work, then stays accountable for the output.

02 / Honest reservation

Intensity needs codification.

The audit shows reusable skills and documented commands lagging far behind usage. The next level is to turn personal operating instinct into shared playbooks, evaluation sets, and measurable standards.

03 / Recommendation

Strong fit for ownership.

I would recommend him for a role where knowledge quality is both editorial and operational—especially when the mandate is to build the system, not merely maintain a queue.

Assessment scope: Claude's published local audit, the application materials available in this workspace, and observed collaboration on this build. It is not an employment reference or a claim of access to private platform analytics.

First 90 days

Start with truth. End with a learning system.

1

Days 01–30

Establish truth

Map the product and policy surface, identify owners, baseline the current agent, and create the first 100-question evaluation set.

2

Days 31–60

Build the quality engine

Launch the gap register, structured documentation pattern, failure taxonomy, release gates, and a weekly review rhythm.

3

Days 61–90

Make it compound

Expand systematic coverage, automate regression testing, publish the quality dashboard, and prove improvement against the baseline.

The close

I know how to learn fast.
I want to build the system that helps the agent do the same.