Skip to main content
AI architecture

AI architecture, deployed to production by a senior engineer

An AI architecture consultant designs how an AI system fits into your real stack and takes it to production: which model and pattern to use, how it reads and writes your data, how it is secured, monitored and kept affordable, and what happens when the model is wrong. It is for business owners and CTOs who have a demo, a pilot or a clear use case and need it to run every day without surprises. I am Francisco Salazar, an independent AI systems engineer and Sr DevOps Engineer with 10+ years of production infrastructure and CKA + CKAD certifications, so the AI is engineered like infrastructure, not glued together in a no-code tool. You work directly with the engineer who builds. A scoped build ships in a three-week sprint for a single fixed price agreed up front, runs in your own accounts, and carries a 60-day operational guarantee after handoff.

What does an AI architecture consultant actually decide?

An AI architecture consultant decides four things before any code is written: which pattern solves the problem (a single prompt, retrieval, an agent or fine-tuning), which model runs each step, how the system reads and writes your company data, and what it does when it fails. Those decisions set the cost, the accuracy and the maintenance load of the system for years.

The work starts from the business process, not from the model: which step eats the time today, what a wrong answer costs, who has to approve the output, and where the data lives. Only after that does the model choice matter. I build with the Claude API, the OpenAI API and Google Gemini, pick per task and budget, and design the system so a provider can be swapped without a rewrite.

I also do not resell tools. If the right answer is a 40-line script or a feature your CRM already has, I say so on the first call.

  • Pattern and model selection grounded in your actual data, volume and budget
  • Data design: what the system may read, what it may write, and under whose permissions
  • Retrieval and memory design so answers cite real sources
  • Failure design: what happens when the model is wrong, slow or unavailable
  • Deployment on your stack or cloud (GCP, AWS, Kubernetes) with a clean handoff

Single prompt, RAG, agents or fine-tuning: which pattern fits your problem?

Most business problems need the simplest pattern that works. A single prompt fits tasks where everything the model needs arrives in the request. RAG fits questions that depend on your documents. Agents fit multi-step work that requires calling tools and making decisions along the way. Fine-tuning fits a narrow, stable task with many clean examples, and it is rarely the first move.

Real systems mix these patterns. An intake flow can classify a message with a single prompt, answer from documents with retrieval, and hand only the cases that need a CRM lookup or a calendar check to an agent. The rule I follow: start with the simplest pattern, measure it against real examples, and move up only when the evaluation shows the simpler option fails.

  • Single prompt: classification, extraction, summaries and drafts from an email or a form. Cheapest to run and easiest to test. Wrong when the answer depends on information the model has never seen.
  • RAG (retrieval-augmented generation): the system searches your documents, passes the relevant passages to the model and answers citing them. Right for policies, manuals, catalogs and contracts that change. Most of the engineering is in the retrieval (chunking, permissions, freshness), not in the prompt.
  • Agents: the model decides which tool to call next, such as looking up a customer, checking a calendar or drafting a reply. Right when the path varies case by case. Agents cost more per task and fail in more ways, so they need step limits, narrow tool permissions and human approval before anything irreversible.
  • Fine-tuning: training a model on your own examples. Right for a consistent format or tone at high volume when you have clean labeled data. Wrong as a way to teach a model your facts, because facts change and retrieval can be updated the same day.

What does production readiness mean for an AI system?

Production readiness means the system can be left running without someone watching it: credentials are stored safely, spending has a ceiling, quality is measured against real examples, every request leaves a trace, a failure degrades to something safe, and a person approves anything that cannot be undone. A demo has none of this and looks identical on a screen.

This is where a production infrastructure background changes the result. A model call is an unreliable network dependency with a variable bill, and I treat it like one: versioned prompts and workflows, backups, error handling, and logs someone can read at the moment something breaks. None of it shows in a demo, and all of it decides whether your team still trusts the system a few months in.

  • Secrets: API keys and credentials live in a secrets manager or the platform vault, never in a prompt, a workflow note or a repository. Each integration gets the narrowest permission that works.
  • Cost guardrails: spending limits at the provider, a budget per request, caps on agent steps, caching where answers repeat, and the cheaper model wherever the evaluation shows it is good enough.
  • Evaluation: a set of real examples with expected answers, run before every prompt or model change, so quality is a measurement and not an opinion.
  • Observability: every request logged with its inputs, retrieved sources, output, latency and cost, plus alerts when errors or spending leave the normal range.
  • Fallback: retries with backoff, a second provider or a simpler path when a model is down, and a safe default such as handing the case to a person.
  • Human approval: anything irreversible (sending, invoicing, deleting) waits for a person. The system drafts, a human confirms.

What does an AI build look like from first call to handoff?

A typical build has six parts: a trigger (an email, a form, a WhatsApp message, a schedule), a data layer that fetches what the model needs, one or more model calls, guardrails around them, a human approval step, and a deployment in your own accounts. A scoped build ships in a three-week sprint, and larger ones take three to five weeks.

Take an inbound request that has to be answered from your documents. The trigger receives the message. A retrieval step searches a Postgres database with pgvector, often on Supabase, that holds your indexed documents, filtered by what that user is allowed to see. A model drafts the answer and cites the passages it used. Guardrails check that the answer has sources and stays inside the cost limit, and route the case to a person when retrieval finds nothing relevant. If the reply goes to a customer, it waits in an approval queue.

Where it runs depends on what you already have. Orchestration lives in n8n or in a Python and FastAPI service, depending on complexity. Hosting is Railway or Vercel for a small footprint, or your own GCP, AWS or Kubernetes environment if you run one. Everything is created in your accounts with your API keys.

  • You own 100% of the code, workflows, prompts and data
  • Exported code and workflows in your own repository
  • Documentation plus Loom walkthrough videos, so the system survives the "hit by a bus" problem
  • 60-day operational guarantee after handoff, and optional monthly support after that

How is a fractional AI engineer different from a no-code agency or a full-time hire?

A fractional AI engineer is a senior engineer you engage for a scoped build instead of a salary: the same person designs, builds and deploys the system for a fixed price, and leaves you owning everything. A no-code agency sells workflows inside a visual tool. A full-time AI engineer makes sense when AI work is continuous and central to your product.

No-code tools are a good call for a proof of concept and for simple automations, and I run n8n in my own operation every day. The limits appear with volume, edge cases, compliance, and anything that needs real code, custom retrieval or deployment inside your cloud. A senior engineer uses those tools when they fit and drops down to Python when they do not.

A full-time hire is the right move when you have a multi-year AI roadmap and enough work to keep a specialist busy. For most small and mid-sized companies the reality is one or two systems that must be designed well once and then maintained lightly. Recruiting a senior profile is slow, the salary runs all year, and one hire rarely covers both the AI layer and the infrastructure under it.

The trade-off is real: I am a solo practitioner. I bring in a second engineer for large scopes, but I am not a team of ten, and a program that needs one should hire one.

What does an AI architecture consultant cost?

I do not bill hourly for builds. Every engagement is a single fixed price agreed up front after a discovery call, with a closed scope. The public reference points are the three productized systems on this site: the Lead Acquisition Engine at USD 4,000 setup, the AI Intake Agent with knowledge base at USD 4,900, and the AI Operating System at USD 6,800.

Each has an optional monthly for operations and iteration: USD 490, USD 1,200 and USD 1,800 per month respectively. A custom architecture is quoted the same way, as a fixed price after discovery. What moves the price is scope: how many systems the AI has to read and write, how messy the data is, how much evaluation the risk demands, and whether it deploys to a managed platform or into your own cluster.

Running costs are separate and stay in your name. Model API usage and hosting are billed to your own accounts, which is also what makes the cost guardrails meaningful. The build carries a 60-day operational guarantee: if something I built breaks in that window, it is fixed at no cost.

When do you NOT need an AI architect?

You do not need an AI architect when the task lives inside one tool, touches no sensitive data and a wrong answer costs nothing. Using ChatGPT or Claude to draft emails, turning on the AI feature your CRM already includes, or connecting two apps with a simple automation all work fine without architecture. Buy the subscription, write a good prompt and move on.

You also do not need one yet if the process itself is undefined. AI does not fix a workflow nobody can describe, and the honest advice in that case is to write the process down first and come back.

Architecture starts to pay when the system touches real company data (documents, CRM, ERP), serves customers or several teams, has a cost per request that matters, or can cause damage when the model is wrong. If none of those is true, save the money.

Frequently asked questions

Do you only build with one AI provider?

No. The architecture picks the model that fits the task and budget, and is designed so a provider can be swapped without a rewrite. I work with the Claude API, the OpenAI API and Google Gemini. Production AI should not be locked to a single vendor, because prices, limits and model quality change faster than your business process does.

Can you deploy into our existing cloud or Kubernetes cluster?

Yes. With CKA and CKAD certifications and 10+ years of production DevOps, deploying into your GCP, AWS or Kubernetes environment, with proper secrets management and observability, is the core competency. If you do not run a cluster, a managed platform such as Railway or Vercel is usually the simpler and cheaper choice, and I will say so.

What is a fractional AI engineer?

A senior engineer engaged for a defined scope instead of hired on payroll. You get the architecture decisions, the build and the deployment from one person, for a fixed price, and you own the result. It fits companies that need one or two AI systems done properly, not a permanent AI team.

How is this priced?

A single fixed price agreed up front after a discovery call, not hourly. A scoped build takes three weeks, larger ones three to five, with a 60-day operational guarantee after handoff. The productized systems on this site, from USD 4,000 to USD 6,800, are the public reference points.

Who owns the code, the prompts and the data?

You do, 100%. Everything runs in your own accounts and infrastructure: your cloud, your n8n, your API keys. Handoff includes documentation, the exported code and workflows in your repository, and Loom walkthrough videos, so another engineer can take over without me.

When are you not the right choice?

When you need a large team, a program that runs for years, or a permanent in-house AI function. I am a solo practitioner who brings in a second engineer for large scopes. I am also not the right choice if the problem is solved by a subscription or a feature you already pay for, and I will tell you that on the first call.

Do you work with companies outside Chile?

Yes. I am based in Chile and work remotely with companies in the US, Canada and Europe, in English or Spanish, with full overlap with US business hours.

Let's talk for 15 minutes about your operation.

A free call. I'll tell you straight which processes today's AI can solve and what your infrastructure needs for them to actually run. If it fits, we move forward. If not, I point you the right way, free.