AI agent development services: agents that run inside your business, not in a demo
AI agent development services cover the design, build and deployment of software agents that use a language model to read a situation, decide what to do, call your systems to do it, and know when to hand the case to a person. This page is for owners and operations leads at small and mid-sized companies who want an agent working inside their business, not a chatbot demo. I am Francisco Salazar, an independent AI systems engineer and Sr DevOps Engineer with ten years of production infrastructure (CKA, CKAD). I build agents on the Claude and OpenAI APIs with tool calling against your CRM, inbox, documents and databases, run them on n8n or a Python service in your own accounts, and put human approval in front of anything irreversible. Scoped builds ship in three weeks at a single fixed price, with a 60-day operational guarantee, and you own all of the code, prompts and data.
What is an AI agent, and how is it different from a chatbot or a workflow?
A workflow is a fixed sequence: when a form arrives, do step A, then B, then C. It is predictable and cheap, and it cannot handle a case nobody drew on the diagram. A chatbot answers questions in a chat window; the good ones read your documents before answering, but they stop at the reply. An agent is different in four ways. It has tools: it can search your CRM, read an email thread, query a database, create a ticket or draft a quote. It has memory: it keeps the context of the case across steps and, when needed, across days. It makes decisions: given the goal and what it found, it chooses the next action instead of following a branch someone hard-coded. And it runs in a loop: act, observe the result, decide again, until the job is done or a stopping rule fires.
That loop is what makes agents useful and what makes them dangerous when built carelessly. A workflow that fails, fails visibly at a step. An agent that is given the wrong tools or no stopping rules can do the wrong thing confidently and repeatedly. Most of the engineering in AI agent development is not the prompt; it is deciding which tools the agent gets, which actions require a person, how the agent knows it is done, and how you see what it did.
- Workflow: fixed steps, no judgment, cheap and predictable
- Chatbot: answers in a window, reads documents, does not act
- Agent: tools, memory, decisions, and a loop with stopping rules
- Rule of thumb: if every case follows the same path, build a workflow; if a person today has to read, look things up and decide, an agent may fit
Which AI agents actually pay off for a small or mid-sized business?
After enough discovery calls a pattern shows up. The agents that earn their cost do one narrow job that a person today does many times a day, with a clear definition of done and a clear place to escalate. The ones that disappoint are the "assistant for everything" that impresses in a meeting and that nobody trusts with a real customer.
Five types come up most often, and they are the ones I quote with confidence:
- Intake and qualification: reads every inbound lead or request from the web form, email or WhatsApp, asks the missing questions, checks it against your criteria, and creates the deal with a summary, or politely closes the ones that do not fit
- Support triage with escalation: classifies each ticket, answers the ones covered by your knowledge base, and routes the rest to the right person with the context already gathered
- Document processing: extracts fields from invoices, purchase orders, contracts or forms, validates them against your records, and pushes the result into your ERP or sheet with exceptions flagged for review
- Research and enrichment: takes a company or a contact, gathers what is public, and writes the structured record your sales team needs before a call
- Internal ops copilot: a chat interface over your own systems (CRM, tickets, calendar, documents) that answers "what is the status of X" and drafts the next action, with write access limited to what you approve
How do I build AI agents, and on what stack?
The model is the smallest part of the decision. I build on the Claude API (Anthropic) and the OpenAI API, and use Google Gemini where it fits, chosen per task on quality, cost and latency rather than loyalty. The agent's tools are defined as functions that call your actual APIs: the CRM, the inbox through the Gmail or Google Workspace APIs, WhatsApp through Meta's Cloud API, your ERP, your database. The agent never gets a generic "browse the internet" tool when a precise "look up customer by email" tool will do; narrow tools are safer and cheaper.
The runtime is either n8n, self-hosted or cloud, when the agent lives naturally among other workflows and your team wants to see the flow, or a Python service with FastAPI when the logic is heavier, needs tests, or has to handle volume. Memory and retrieval live in Postgres with pgvector, on Supabase or your own instance, so the agent can remember the case and search your documents with the same database you already back up. It runs on Railway, Vercel, GCP or AWS in your accounts, containerized, with secrets managed properly.
Two things are non-negotiable in every build. First, human approval before anything irreversible: sending to a customer, invoicing, deleting, changing a record that other systems depend on. Second, evals and logging: a fixed set of test cases the agent must pass before any prompt change goes live, and a log of every run with the inputs, the tools called and the output, so you can audit what it did and I can improve it with evidence instead of guesses. I run Claude-based agents in my own operation every day on exactly this setup.
What does an AI agent build look like, step by step?
Take the most common request: an intake agent for a company where every inquiry that arrives through the web form, a shared inbox and WhatsApp is read by the owner or by one salesperson. Here is what the build looks like in plain words.
- Trigger: a new form submission, an email in the sales inbox or a WhatsApp message on the business number starts a run. Each channel is a webhook into the runtime; nothing polls.
- Data: the agent receives the message, the sender's history from the CRM, and a retrieval step over your service descriptions, pricing rules and FAQs stored in pgvector. It does not see anything it does not need.
- Model and tools: Claude or GPT with five tools: look up contact, search knowledge base, ask the prospect a question, create or update the deal, hand off to a human. The prompt is versioned in your repository like code.
- Guardrails: the agent can ask at most a fixed number of questions, cannot quote a price outside the rules you gave it, and stops on any doubt. Every run is logged with the tools it called.
- Human approval: qualified leads are created in the CRM with a summary and a proposed reply; the reply goes out only when your salesperson approves it in Slack or in the CRM. Clear disqualifications close politely on their own; anything uncertain waits for a person.
- Where it runs and what you own: n8n or a FastAPI service in your Railway or cloud account, Postgres in your Supabase project, API keys in your name. Handoff includes the code and workflows in your repository, documentation, and Loom walkthroughs.
What do AI agent development services cost?
I do not bill hourly for builds. After a free discovery call you get a written proposal with a closed scope and a single fixed price, delivered in three weeks for a scoped agent and three to five weeks for larger ones. The reference points are the two productized systems published with list prices on this site: the AI Intake Agent with knowledge base at USD 4,900 (optional USD 1,200 per month for operations and iteration), and the AI Operating System at USD 6,800 (optional USD 1,800 per month). A custom agent that is close to one of these is priced the same way; a scope that is much smaller or much larger is quoted as a fixed price after discovery, never as an open-ended hourly engagement.
Infrastructure and API subscriptions (LLM usage, hosting, n8n Cloud if you choose it) are billed to you directly, in your accounts, with a cap declared in the proposal. Because every run is logged, you can see the exact model cost per case instead of guessing. The 60-day operational guarantee covers anything I built that breaks after handoff, at no cost.
Why do AI agents bought from a demo fail in production?
A demo agent is optimized for the meeting: it answers the five questions the seller rehearsed, on the seller's data, with no consequences if it is wrong. Four failure modes show up when the same thing meets real customers and real records.
- No guardrails: the agent can send, promise or change things nobody authorized it to, and the first bad case ends up in a customer's inbox
- No observability: nobody can say what the agent did last Tuesday, which tool it called or why, so a failure is discovered by the customer and cannot be reproduced
- Vendor lock: the prompts, the memory and the integrations live in someone else's platform, so leaving means rebuilding from zero and staying means paying whatever the price becomes
- No evals: every prompt tweak is tested by hand on two examples, so fixing one case silently breaks another
When should you not build an AI agent?
An agent is the wrong tool more often than sellers admit. If every case follows the same steps, a workflow in n8n does the job for a fraction of the cost and with none of the ambiguity. If the volume is a handful of cases per week, a person with a good template is cheaper than any system. If the data the agent would need is not in any system yet, because it lives in someone's head or in unstructured chats, the first project is getting it into a CRM or a database, not the agent. And if your CRM already has the feature, use it. I do not resell tools, and I will say any of this on the first call.
I also decline scopes where the agent would take irreversible actions with no human in the loop from day one, or where the promise is "replace the team". Those builds fail, and the reputational cost lands on you. What works is an agent that removes the reading, the looking-up and the first draft, and leaves the judgment where it belongs until the logs show it can be trusted with more.
How does the engagement work?
It starts with a free 15-minute call to see whether an agent is the right tool and whether I am the right person. If there is a fit, you receive a written proposal with closed scope, timeline and one fixed price. During the build you get weekly updates and approve before anything touches production or a customer. Handoff includes documentation, the exported code and workflows in your repository, Loom walkthroughs so the system survives staff changes, and the 60-day guarantee. Optional monthly support after that covers iteration, new tools for the agent and monitoring.
I work remotely from Chile with companies in the US, Canada and Europe, in English or Spanish, with full overlap with US business hours. It is a solo practice; for large scopes I bring in a second engineer, and you always work directly with the person who builds.
Related
- Custom AI developmentThe hub: every kind of custom AI build, how I scope them and what they cost.
- AI knowledge base chatbotWhen the job is answering from your documents, not acting on systems.
- WhatsApp AI agent for businessThe same agent pattern on the channel your customers already use.
- n8n consultantWhen the agent runs on n8n workflows: builds, rescues and self-hosting.
- AI Intake AgentThe productized intake and qualification agent, USD 4,900 list price.
- AI Operating SystemThe ops copilot plus workflows, USD 6,800 list price.
Frequently asked questions
Do you build AI agents with Claude or with OpenAI?
Both, and Google Gemini where it fits. The model is chosen per task on quality, cost and latency, and the agent is built so the model can be swapped without rewriting the tools or the memory. I am not tied to a vendor and I do not resell any of them.
Can the agent connect to my CRM, ERP, inbox or WhatsApp?
Yes, through their APIs. HubSpot and Pipedrive-style CRMs, Gmail and Google Workspace, WhatsApp via Meta's Cloud API or a business solution provider, Slack, Cal.com, Postgres and Supabase, and most ERPs with an API. If a system has no API, we discuss a safe workaround on the call before quoting.
Can you build an agent that acts fully on its own, with no human approval?
Not for irreversible actions on day one, and I will say so. Sending to customers, invoicing and deleting always start behind an approval step. Once the logs show the agent is reliable on a class of cases, we can widen its autonomy deliberately, with evals guarding each change.
What do you not do?
I do not sell generic chatbots from a template, I am not an n8n certified partner or a reseller of any platform, and I do not take scopes whose premise is replacing a team. If the right answer is a 40-line script or a feature your CRM already has, you hear that on the first call and there is no project.
Who owns the agent after handoff?
You do, 100%. The code, workflows, prompts, memory and data live in your repository and your accounts, with API keys in your name. Handoff includes documentation and Loom walkthroughs, so there is no dependency on me after the 60-day guarantee unless you choose the optional monthly support.
How long does an AI agent build take?
Three weeks for a scoped agent such as intake, triage or document processing: one week of discovery and design with your real data, one of build against your systems, one of evals with your team and handoff. Larger scopes, such as a copilot over several systems, take three to five weeks.
Do you work with companies outside Chile?
Yes. I work remotely from Chile with companies in the US, Canada and Europe, in English or Spanish, with full overlap with US business hours, and I invoice in USD. Discovery, weekly updates and handoff happen over video, and the systems run in your accounts, so location never touches the delivery.
Bring the process a person reads and decides on every day. We will see in 15 minutes whether an agent fits.
A free call. I'll tell you straight which processes today's AI can solve and what your infrastructure needs for them to actually run. If it fits, we move forward. If not, I point you the right way, free.
