For the past few weeks I've been chewing on one question: how do I build an AI harness around TalonHR using the MCP server we already have?
On paper it sounds easy. Give a language model access to the tools TalonHR exposes, let it work out what the user wants, and let it act on their behalf.
Then you start building it, and software does what software always does: it takes a simple idea and finds a dozen ways to make it complicated.
I'm still in the middle of the architecture. Some pieces already exist in TalonHR; others have to live in the harness. Most of my thinking time has gone into tool discovery, execution, authentication, context, and the unglamorous question of how to make the whole thing reliable enough that anyone would actually want to use it.
So here's the approach I'm taking, the decisions behind it, and where I want it to go.
First, what am I actually building?
Not another chatbot. The world has plenty of those.
What I'm building is a harness: a layer of software that sits between a language model and the things TalonHR can already do.
A language model is pretty good at understanding what you're asking. What it doesn't have is any idea what's going on inside your application. It doesn't know which jobs are open, who has applied, or which interviews still need a time slot.
For that, it needs tools. And handing a model some tools is only half the job. Something still has to manage those tools, keep track of the conversation, check what the model is trying to do, deal with failures, and stop it from doing things it shouldn't.
That something is the harness.
The goal is to let people talk to TalonHR in plain language, while the real business logic stays inside the application, where it belongs.
Say an employer wants to see who's waiting for an interview. Today that means clicking through several screens. With the harness, they could just ask:
"Which candidates are waiting for interviews this week?"
The harness works out what they mean, picks the right tools, pulls the data from TalonHR and shows it in a way that's easy to read.
The key point is that the AI isn't making up the answer. It's fetching real data through the same interfaces the application already uses.
The starting point: our MCP server
The good news is I'm not starting from zero.
TalonHR already has an MCP server, documented at https://talonhr.eu/docs/mcp.
MCP (Model Context Protocol) is a standard way for AI applications to find and use external tools. Think of it as a menu the model can read.
We split TalonHR's MCP into two endpoints.
The public one lives at:
https://mcp.talonhr.eu/public/mcp
It handles the things anyone can do: searching for jobs, getting job details and reading public resources.
The authenticated one lives at:
https://mcp.talonhr.eu/mcp
What you can do here depends on who you are and which workspace you're in. It covers applicant, employer and recruiter workflows.
That split matters. The harness shouldn't have a master key to all of TalonHR. It should have exactly the same access as the person using it, and no more.
A recruiter shouldn't be able to see another employer's private candidates just because they asked the AI nicely. That would be a very creative way to ship a security hole.
The architecture
I'm leaning towards Go for this.
Partly because it's what I use most, and partly because I want the harness to be a small, predictable service, not a pile of dependencies held together by hope.
The design splits the work into a few clear jobs:
User
|
v
Conversational UI
|
v
AI Harness (Go)
|
+--------+---------+
| |
LLM Client Context Manager
|
Tool Orchestrator
|
Policy / Validation
|
MCP Client
|
+---+------------------+
| |
Public MCP Authenticated MCP
| |
+----------+-----------+
|
TalonHR
The harness sits in the middle. It runs the conversation and coordinates between the model and TalonHR.
The model figures out what the user is asking and picks the tools that might help. The MCP client talks to TalonHR. The policy layer checks what the model is trying to do before anything happens. And the context manager remembers the conversation, the current workspace and what's already been fetched.
Keeping these separate is deliberate. I want to be able to swap the language model without touching the MCP code, and change the MCP side without rebuilding the whole conversation system.
Tool discovery comes first
One of the first pieces is working out how the harness finds and manages the tools on offer.
MCP has a tools/list operation that returns every tool a server exposes. TalonHR's public server, for example, offers tools like search_jobs and get_job.
So rather than hand-teaching the model every endpoint, the harness can discover the tools, read their descriptions and schemas, and translate them into whatever format the model expects.
Roughly, it goes like this:
Connect to MCP server
|
v
Discover available tools
|
v
Read tool descriptions and schemas
|
v
Register permitted tools with LLM
|
v
Wait for user requests
There's a design choice hiding in there. Just because a tool exists doesn't mean the model should see it.
A public job-search assistant has no business knowing that employer interview-scheduling tools exist. That's like handing a visitor at reception the keys to every office because they happen to be on the same keyring.
So I treat discovery and availability as two separate jobs. The MCP server tells the harness what exists. The harness decides what's allowed right now.
That matters even more once logged-in workflows come into play.
The execution loop is where it gets fun
A traditional chatbot takes a message and sends back a reply. Done.
An agent has more work to do. It might call a tool, look at the result, call another tool, and only then answer.
Say someone asks:
"Find remote Go backend jobs."
The model works out that it needs a job search and asks for the search_jobs tool. The harness checks the arguments, calls the MCP server and hands the results back. The model then turns those results into something a human would actually want to read.
A stripped-down version of the loop in Go looks like this:
func Run(ctx context.Context, messages []Message) error {
for step := 0; step < maxSteps; step++ {
response, err := llm.Generate(ctx, messages, tools)
if err != nil {
return err
}
if response.IsFinal() {
return sendResponse(response)
}
for _, call := range response.ToolCalls {
if err := validate(call); err != nil {
return err
}
result, err := mcp.CallTool(
ctx,
call.Name,
call.Arguments,
)
if err != nil {
messages = appendToolError(messages, call, err)
continue
}
messages = appendToolResult(messages, call, result)
}
}
return ErrMaxStepsReached
}
This is a sketch of the mechanism, not the finished thing. The real version has to deal with each model's own tool-call format, timeouts, cancellation, authentication, and the different kinds of responses MCP tools send back.
One thing I'm watching closely is how long the agent is allowed to keep going. Left alone, a model can happily call tools over and over, retry things that keep failing, and burn through tokens without getting anywhere. Think of an intern who never asks for help, except you're paying for every word they say.
So step limits, execution budgets and clear stopping rules are part of the core design, not something I'll bolt on later.
Authentication is the tricky part
Public tools are fairly easy. Logged-in workflows are a different story.
TalonHR has different personas: applicants, recruiters and employers. One person can also belong to several workspaces, each with its own permissions. The harness has to keep all of that straight.
If an employer asks to schedule an interview, the system needs to know which workspace they're in, which candidate they mean, and whether they're allowed to do it at all.
The language model should not be the one making those calls, and I'm very deliberate about this. The model can suggest an action. The harness checks it. TalonHR has the final say on who can do what.
I'm also working out how to handle OAuth sessions and token refresh without the model ever seeing a credential. The model should work with what the user is allowed to do, not with keys that open the whole application.
Then there's context
This one has kept me busy.
Say a recruiter asks:
"Show me the candidates in my talent pool."
The harness fetches them. Then the recruiter asks:
"Which of them have experience with Go?"
And then:
"Show me the three strongest matches."
To a person, those questions obviously belong together. The system needs to understand that too.
But you can't just stuff the whole conversation, plus every tool result, back into the model each time. That gets expensive fast. It's also a privacy problem when you're dealing with candidate data.
The approach I'm exploring is for the harness to keep a structured record of the conversation, while the model gets a smaller working version. Instead of re-sending full candidate profiles every turn, the harness holds on to IDs, filters and references to what's already been fetched. The model gets enough to follow along, and the harness can pull more detail when it's needed.
One more thing: I don't want a model's ranking of candidates to be treated as an objective hiring verdict. "Strongest match" needs clear criteria, evidence behind it and a human checking it.
The hard part is the balance. Give the model enough to be useful, but not so much that it's lugging around everything it has ever seen.
What happens when the AI wants to change something?
Reading data is one thing. Changing it is another.
TalonHR's MCP server already asks for confirmation before protected write operations, and I plan to keep it that way.
Take interview scheduling. The AI might find a good slot and prepare the booking, but nothing happens until the user approves it:
User requests interview
|
v
AI identifies scheduling tool
|
v
Harness validates arguments
|
v
MCP prepares confirmation
|
v
User reviews and approves
|
v
MCP executes operation
|
v
Harness reports result
I don't want the harness to become a back door around TalonHR's security. The same rule applies to job applications, moving candidates between stages, and anything else with real consequences.
The AI can prepare and help. The application enforces the rules.
Planning for failure early
If years of backend work teach you anything, it's that the happy path is the boring part. The real questions show up when something breaks.
What if the model calls a tool with bad arguments?
What if the user's login expires halfway through a task?
What if an MCP call times out after TalonHR has already done the work?
What if the model keeps picking the same tool and it keeps failing?
I want every tool call to be traceable, every failure to make sense, and every retry to be under control.
For write operations, I especially don't want blind retries when we don't know what happened. An interview booking that timed out isn't necessarily an interview booking that failed. Retry it carelessly and a candidate gets invited to the same interview twice, which is a great way to look keen and a terrible way to look organised.
Then there's trust
A harness doesn't only take instructions from the user. It also reads whatever the tools send back.
For TalonHR, that means job descriptions, candidate profiles and other content written by users. Some of it could be written specifically to manipulate the model.
Picture a job description with a line slipped in telling the assistant to ignore its instructions and go fetch something unrelated. Not quite the "requirements" section you were hoping for.
The harness has to treat that content as data, not as orders.
I'm planning a strict separation between system instructions, user requests and tool responses, with checks at the point where tools actually run. The model can read and interpret outside content, but that content should never be able to change the agent's permissions or rules.
A clever system prompt isn't enough here. The security boundary has to live in code.
Where things stand
The project is still taking shape. The MCP infrastructure gives me a solid base, and the core design is getting clearer.
What I'm focused on now is fitting the pieces together without building something more complicated than it needs to be.
Next up are the MCP client, tool discovery, the execution loop and the context approach. After that come logged-in workflows, approvals and proper testing.
For testing, I want to use real recruitment tasks, not tidy examples designed to make the model look clever. A demo is easy to nail when you write the questions yourself. Handling vague requests, missing information, permission errors and users doing unexpected things is much harder.
That's where the real engineering is.
Will this replace parts of the interface?
Maybe.
One thing we've learned building TalonHR is that more features don't automatically make software easier to use. Often it's the reverse. Every new feature brings another screen, another menu, another workflow to learn.
An AI interface offers a different route. Instead of needing to know where a feature lives, you just say what you're trying to do, and the harness turns that into actions the application supports.
I don't think that kills the traditional interface. Plenty of tasks are still better done visually, especially comparing information, reviewing detailed records or making big decisions.
But I can see the interface shifting from "find the right menu" to "say what you want".
One last thought
The deeper I get into this, the more convinced I am that the language model is only one part of the problem.
The real work is everything around it: tool execution, authentication, authorisation, state, context, observability, cost control and failing in predictable ways.
None of that is new. We've been wrestling with these problems in distributed systems for years. What's new is adding a component that reads plain language and decides for itself which operations to run. That opens up some great possibilities, and some fresh ways for things to go wrong.
For now, I'm keeping the scope small on purpose. I'd rather have a harness that does a handful of useful things reliably than one that claims to run an entire recruitment platform and falls over the first time someone asks something unexpected.
There's plenty left to build, and I'm sure some of these decisions will change once I'm deeper in. That's half the fun.
I'll share more as it comes together, especially once tool execution and logged-in workflows are working together.
If you want to dig into the integration yourself, the TalonHR MCP docs are here: https://talonhr.eu/docs/mcp