AI Agents with Human Approval

Agents that ask before they act.

AI agents that do the legwork (reading, drafting, classifying, planning) while deterministic code owns money, stock and dates, and a person approves anything consequential. Every run leaves an evidence trail you can read afterwards.

The problem

You want the speed, not the risk.

The pitch for AI agents is simple: let the model handle the repetitive work. Read the incoming request, work out what is being asked, draft the reply, update the system. And for a lot of that work, models are genuinely good.

The trouble starts at the edges. A model that invents a price. A reply sent before anyone read it. A scarce resource handed out because the request sounded urgent. A confident answer built on last month's instruction, which someone replaced last week. In a demo, these are funny. In your business, they cost money and trust.

So the real question isn't whether an agent can do the work. It is how to let it do the routine parts on its own while it stops at the parts that need a person.

What we build

Agents with guard-rails built in.

  • Intake and quoting agents. The model extracts requirements and spots missing information; code calculates the price and availability; a person approves before anything is sent.
  • Scope and document review. Scattered briefs, notes and emails compared against an agreed baseline, with every finding citing its source, and a person approving each requirement.
  • Operations agents. An agent that works through narrow tools (classify, check stock, reserve, schedule, draft a notification) and escalates anything that breaches a protected threshold.
  • Agents with memory. Decisions, instructions and commitments stored as typed records with sources and timestamps, so the agent can say what is current and why.
  • Risk-gated decision systems. Several agents propose, a critic objects, and a deterministic governor has the final say, with "no action" treated as a valid, audited result.
  • AI inside developer tools. Deterministic checks decide what is risky; the model proposes a fix; the same checks judge the fix.
How we work

Scope, build, hand over.

1. Map the process and mark the judgement calls. We list each step and sort it: routine, needs a rule, or needs a person. That map is the design.

2. Put the rules in code, not the prompt. Prices, limits, thresholds and allow-lists live in tested functions and the database. The model reasons; code enforces.

3. Validate everything that crosses a boundary. Model outputs are checked against a schema before anything acts on them. Unknown citations are rejected. Missing information pauses the flow.

4. Build the approval screen properly. The person approving sees what the agent saw, what it proposes and why, and can approve, reject or edit, with the decision recorded.

5. Deploy and hand over. Keys stay server-side, health checks never expose credentials, and you get a runbook covering the model, the tools and how to switch the agent off.

Proof

Seven builds, one pattern.

These are team hackathon builds and prototypes, not client deployments, and each project page credits the people who built it.

  • QuotePilot AI (Qwen Cloud hackathon). Qwen analyses the request; deterministic code sets price and availability; a mandatory human approval checkpoint sits before the simulated send.
  • ScopeGuard AI (OpenAI Build Week). Strict schema output where every finding cites a source ID, and final confirmation is blocked while questions are open.
  • NeighborOps AI (Agents for Humans hackathon). A Strands agent with 15 narrow tools, no general database access, and a protected-reserve threshold that escalates to a person.
  • ThesisCircuit (Alpaca AI Trading Agents hackathon, paper trading only). Competing strategy agents, a critic, and a fail-closed risk governor no agent can override.
  • ClientOps Memory AI (CockroachDB × AWS hackathon). Typed, evidence-backed memory on CockroachDB with Amazon Bedrock; its evaluation harness passes 10 of 10 long-term-memory scenarios.
  • LockSmith (IBM Bob 2.0 hackathon, built solo by our founder). Twelve deterministic rules decide what is risky; IBM Bob rewrites the migration; the same gate re-checks the rewrite.
  • CaptionForge AI (AMD Developer Hackathon). Vision first, text fallback, labelled mock last, with the source shown on every result and human review before posting.
Stack

What we reach for.

Next.js and TypeScript or Python and FastAPI; OpenAI, Qwen, Amazon Bedrock, Fireworks or local models; Strands Agents where an agent framework earns its place; Zod or JSON Schema for validation; PostgreSQL or CockroachDB (including vector search) for state and memory; and automated tests around every deterministic tool.

Not a fit

When we'll say no.

  • You want an agent that sends, spends or commits with no human checkpoint on consequential steps. We won't build that.
  • The underlying process isn't defined yet. An agent will automate the confusion; let's define the process first.
  • You need guaranteed accuracy from the model itself. No one can promise that; we design so a wrong answer gets caught.
Problems we fix

If one of these sounds familiar, we should talk.

I need an AI agent that asks before it acts.

What we doThe model drafts and reasons. Deterministic tools own anything that touches money, stock or dates. A human approval checkpoint sits in front of every consequential action, with an evidence trail you can read afterwards.

The chatbot sounds confident and gets the price wrong.

What we doThe model never sets a price. It picks from an allow-listed catalogue, and code does the arithmetic. If information is missing, the agent asks instead of guessing.

Every AI session starts from zero.

What we doTyped, evidence-backed memory: decisions, instructions and commitments stored with their sources, and superseded decisions kept visible rather than silently overwritten.

What you get

A system, not a pile of parts.

Pick the pieces you need. Most projects start small and grow from there.

  1. Human approval checkpoints

    Explicit stop points before anything is sent, allocated, booked or committed, with approvals and rejections recorded.

  2. Deterministic tools for money and data

    Narrow, tested functions for prices, availability, stock and risk limits. The model calls them; it can't override them.

  3. Evidence trail for every run

    What the agent saw, which tools it called, what they returned and who approved what, stored so you can audit a decision later.

Proof

Related work and reading.

Client work is shown anonymised. Hackathon builds link to their public case studies.

Hackathon · lablab.ai · Sep 2026

LockSmith (case study on talalkhawaja.com)

A deploy gate for PostgreSQL migrations: it flags lock-taking statements, proves the locks in embedded Postgres, estimates blocking time, and has IBM Bob rewrite them for zero downtime.

Hackathon · lablab.ai · Sep 2026

ThesisCircuit (case study on talalkhawaja.com)

Paper-only options research agents: three strategies compete, a critic objects, and a fail-closed risk governor has the final say. NO TRADE is a first-class, audited result.

Hackathon · Devpost · Jun 2026

CivicAI Readiness Blueprint (case study on talalkhawaja.com)

Decision support for community AI readiness: transparent local scoring, gaps, scenario comparison, a three-phase responsible-AI roadmap, and an optional AI policy insight.

Finalist — USAII Global AI Hackathon 2026
Hackathon · Devpost · Aug 2026

ClientOps Memory AI (case study on talalkhawaja.com)

An AI operations agent for agencies that turns conversations into typed, evidence-backed memory in CockroachDB, so decisions, instructions and tasks survive the session.

Hackathon · Devpost · Jul 2026

QuotePilot AI (case study on talalkhawaja.com)

An autopilot quoting agent for service businesses: Qwen reads the request, deterministic tools own price and availability, and a human approves before anything is sent.

Hackathon · Devpost · Sep 2026

NeighborOps AI (case study on talalkhawaja.com)

An operations dashboard for community pantries: a Strands agent handles routine coordination through narrow tools, and people decide how scarce resources are used.

Hackathon · Devpost · Jul 2026

ScopeGuard AI (case study on talalkhawaja.com)

Turns scattered client instructions (briefs, notes, emails, chats) into an evidence-backed scope that a person reviews and confirms before any work is committed.

Hackathon · lablab.ai · Jul 2026

CaptionForge AI (case study on talalkhawaja.com)

Turns a short video clip into four caption styles by sampling frames in the browser and calling Fireworks models, with honest fallbacks and human review before anything is posted.

Stack

What we work with.

  • OpenAI API
  • AI agents
FAQ

Straight answers.

Whichever fits the job and your constraints. Our builds have used OpenAI, Qwen on Alibaba Cloud, Amazon Bedrock (Nova), Fireworks-hosted models and local models on AMD hardware. We keep the model behind a narrow interface so it can be swapped.

Only on steps you have explicitly marked as routine, and only through narrow tools. Anything consequential (sending, spending, allocating scarce stock, committing to a date) stops for a person.

Structure. Outputs are validated against a schema, findings must cite real source IDs, missing information is flagged rather than filled in, and business rules live in code, not in the prompt.

Yes, through narrow tools we write for each action, not a general 'do anything' connection. The agent gets exactly the access each step needs.

The builds listed here are hackathon projects and prototypes, and we label them that way. What carries over to client work is the architecture: narrow tools, validation, approval gates and audit trails.

We don't publish fixed prices. We start by mapping the process and marking which steps are routine and which need judgement, then quote the build. Email info@teqprotech.com.

More in AI

Often paired with this.

Start here

AI Agents with Human Approval, unblocked.

Tell us what it’s doing that it shouldn’t (or not doing that it should). The brief form opens with AI Agents with Human Approval pre-selected.