AI agents

AI agents thatwork, reason and act.

We build AI agents that understand context, use your tools and data, and take action across the workflows that matter to your business.

  • AI agents
  • Agentic workflows
  • Tool use
  • RAG & memory
  • Human-in-the-loop
AI
  • KnowledgeYour data, context, insights

  • MemoryLearn and improve

  • ToolsUse your existing systems

  • ActionsGet things done

  • Research
  • Write
  • Analyse
  • Update
Understand|Decide|Act

Autonomous · Connected · Controlled

  • ISO/IEC27001
  • ISO/IEC42001
  • DPA andBAA ready
  • AWS, Azure,Google Cloud
  • All cloudregions

OrbitNexa designs, builds and operates custom AI agents for banking, healthcare and regulated operations, in the client's own AWS, Azure or Google Cloud account. Every agent ships with tracing, human-in-the-loop gates, an eval suite and a tamper-evident audit log, under our ISO/IEC 42001 and ISO/IEC 27001 certified management systems.

Use cases

AI agents for six industries, compliant with each industry's regulations.

Every agent integrates with systems you already run and is placed on the autonomy ladder per action type, a placement you own.

  • Banking, NBFC & lending

    Decisions with an audit trail a regulator can follow.

    KYC extraction and verification · preliminary underwriting with adverse-action reasons · collections voice agent within RBI conduct rules · AML alert narratives

    RBI FREE-AI · FCA Consumer Duty · DPDP

    ClaudeLangGraphMCPAWS
  • Insurance

    Intake, verification and a recommendation, with a human decision.

    First notice of loss intake · document and image extraction · coverage verification · payout recommendation · underwriting triage

    IRDAI · FCA

    GPTLangGraphpgvectorAzure
  • Healthcare & diagnostics

    Clinical documents drafted, coded and routed, signed by a clinician.

    Prior-authorisation preparation · ambient clinical documentation · lab-report drafting and routing · scheduling and triage · claims appeals

    HIPAA · DPDP · ABDM · NHS DTAC

    ClaudeFHIRpgvectorGoogle Cloud
  • Pharma & life sciences

    Regulated intake and assembly with the record GxP expects.

    Pharmacovigilance case intake and coding · regulatory submission assembly · trial-recruitment screening · literature surveillance

    GxP · 21 CFR Part 11

    ClaudeLangGraphPythonAWS
  • E-commerce & retail

    Agents with real access to orders, returns and stock.

    Order and returns resolution against the OMS · merchandising and replenishment · catalogue enrichment · price monitoring

    Scoped tool permissions · dry-run mode for writes

    GPTNode.jsRedisAWS
  • Enterprise operations

    IT, HR, finance and legal copilots wired into the systems teams already use.

    IT helpdesk resolution · onboarding · invoice and PO matching · contract clause review · policy Q&A grounded in the intranet

    Permission-aware retrieval · SIEM export of every call

    ClaudeServiceNowSalesforceAzure

Custom AI Agents

Decision to action.

  • Traceable
  • Gated
  • Evaluated
  • Human in the loopReviewer identity logged.

Our process

From a golden set to an agent in production.

A pilot on real data, never synthetic; a human-in-the-loop validation gate before production; and a runtime you can operate without us.

  1. 01

    Discover

    Discovery & Strategy

    1-2 weeks

    We map your ecosystem, constraints and KPIs before any engineering starts, so every technical decision has a reason on record.

  2. 02

    Architect

    Architecture & Design

    2-3 weeks

    Architects draw the system and designers prototype the interface, both reviewed and signed off before a line of code is written.

  3. 03

    Build

    Agile Development

    4-12 weeks

    Iterative sprints on modern frameworks, with a senior engineer reviewing every pull request before it merges.

  4. 04

    Evaluate

    Quality Assurance

    2-4 weeks

    Automated unit, integration and acceptance tests, plus performance and security checks against real-world scenarios.

  5. 05

    Operate

    Launch & Evolution

    Ongoing

    A zero-downtime deployment, then continuous monitoring and iterative enhancement as your standing technical partner.

How an agent is evaluated.

Tested on your cases before it ships, and re-tested every week after.

  1. Golden setOne hundred to three hundred real cases with expected outcomes, agreed in discovery and owned by you.
  2. Regression gateThe set runs on every prompt, tool or model change. A drop blocks the release.
  3. SamplingA share of live decisions re-scored every week, with drift alerts to a named owner.
  4. Upgrade rehearsalA new model version runs shadow for an agreed period before it decides anything.
  5. ToolingLangfuse or LangSmith for tracing · promptfoo, DeepEval or Ragas for eval suites · Braintrust or Arize Phoenix where you already run them.

Who is on the engagement.

Overlap with UK and Indian working hours every day, a shared channel, a weekly call and a demo against the golden set every two weeks.

  • Engagement leadA senior AI engineer, not an account manager
  • Two to three agent engineersShipping in your repository
  • Integration engineerFor the CRM, core or EHR connectors
  • Eval and governance ownerKeeps the set and the pack current
  • Your named reviewersSign the eval report before production
The full process, phase by phase

Let’s build together

Have an idea?
Let’s make it real.

Tell us about your goals. We’ll help you find the right way forward.

  • Share your idea

    Tell us what you're looking to build.

  • Explore possibilities

    We'll understand your goals and suggest the right approach.

  • Plan the next steps

    Together we define the roadmap.

  • Build what's next

    Turn ideas into real impact.

Let’s discuss
your project

Whether it’s a new product, a platform upgrade or a complex challenge, we’re here to help.

Get in touch

Engagement models

Start with a pilot on your data. Everything after it is optional.

Shapes described by what happens, how long they run and what you hold at the end. Never an amount.

  • Agent Pilot

    One specific use case, built on your data and measured against the golden set.

    TeamEngagement lead, two agent engineers, your reviewers
    Timeline4 – 6 weeks
    Key deliverablesA working agent, an eval report and a written go/no-go
  • Production Agent Build

    A validated pilot taken to production: integrations, tracing, gates, rollback and the governance pack.

    TeamLead, agent and integration engineers, governance owner
    Timeline8 – 16 weeks
    Key deliverablesAn agent in production with every action traceable to a person or a model version
  • Multi-Agent System

    Coordinated agents across functions, with one orchestration layer and one audit trail.

    TeamLead, three to four engineers, architect, governance owner
    Timeline12 – 24 weeks
    Key deliverablesA system your risk team can read end to end, plus agent operations

Start with a readiness review

2 weeks

Use-case clarity, data access, systems with APIs, a governance owner, eval data, risk tier and timeline, read together. You leave with a readiness band, the two or three gaps to close, and a scoped pilot plan.

Book the review

Runs where your data lives.

Your account, your region, your keys. And you can run it without us.

  • Your cloud account

    AWS, Azure or Google Cloud, India region by default; UK or EU region for UK and EU clients.

  • Model routing by data class

    Public data may reach a hosted frontier model; personal or clinical data goes to an in-account model or a BAA/DPA-covered endpoint. The routing table is a deliverable.

  • Keys and logs stay yours

    Bring-your-own-key encryption, logs in your SIEM, sub-processor list published.

  • You can run it without us

    Runbooks, infrastructure as code in your repository, model and prompt versions in your registry, an exit note in the governance pack.

Governance

Five frameworks, one set of controls, and the artefacts each one produces.

Every one of these artefacts is produced on every engagement, not on request. ISO/IEC 42001 is the anchor; we are certified against it.

Framework · control we build · artefact you receive

  • ISO/IEC 42001 — every agent registered with owner, purpose, model version and risk tier → risk assessment, model inventory entry
  • EU AI Act (Arts 9, 12, 13, 14) — tamper-evident action log, HITL gates by risk tier, user-facing disclosure → logging schema, oversight policy, technical-file section
  • UK FCA Consumer Duty, SM&CR — full-interaction tracing, intervention thresholds, sign-off record → pre-deployment outcome assessment, audit export
  • RBI FREE-AI and model risk — eval report from an independent set, explainability surface, grievance hook → validation report, vendor accountability statement
  • HIPAA, DPDP, UK GDPR — region pinning, model routing by data class, BYOK, PHI never to a non-BAA endpoint → data-flow diagram, sub-processor list, DPA/BAA

Ten agentic threats, and the control against each (OWASP 2026)

  • Goal hijack and prompt injection — input classification, instruction and data separated, injection evals in the regression set
  • Tool misuse and excessive agency — scoped tool permissions per agent, allow-lists, dry-run mode for writes, irreversible actions always gated
  • Identity and privilege abuse — each agent runs as its own least-privilege identity; no shared service accounts
  • Memory poisoning — memory writes validated and versioned; retrieval filtered by document permissions
  • Unsafe code execution and insecure inter-agent messages — sandboxed execution, signed and schema-validated A2A payloads
  • Cascading failures, supply chain, rogue agents — circuit breakers, pinned model versions, vetted MCP servers, every action traced, a kill switch

The autonomy ladder.

Four rungs. Every action type is placed on one, and you own the placement.

  1. 01L1 Suggest — the agent drafts; a person acts. Clinical documentation, underwriting notes.
  2. 02L2 Act with approval — the agent prepares the action; a named person approves; reviewer identity logged.
  3. 03L3 Act and notify — reversible actions execute; a person is notified and can revert.
  4. 04L4 Act autonomously — only for low-risk, reversible, high-volume actions inside hard limits.

FAQ

Questions?
We're here to help.

Straight answers for a risk officer, a CTO and a procurement lead. Still have a question? We're just a message away.

Get in touch
  1. 01What is the difference between an AI agent and a chatbot?
    A chatbot answers questions. An agent makes decisions, takes actions in your systems, and is accountable for the outcome. That is why every agent we ship carries tracing, human-in-the-loop gates, content moderation, prompt-injection defences and a rollback path. A chatbot needs none of that; an agent that acts on a loan file or a claim needs all of it.
  2. 02Which models and frameworks do you build on?
    Anthropic Claude as the primary model, the OpenAI GPT-4 family where it fits, and open-source models (Llama, Mistral) for sovereign deployments that cannot call an external API. Orchestration in LangGraph or CrewAI, with custom orchestration where neither fits; tool integration through MCP where a system supports it and function-calling where it does not.
  3. 03How do you keep an agent auditable?
    Every agent action is logged with the inputs it saw, the tool it called and the model version that decided. Human-in-the-loop gates sit in front of any irreversible action. Evals run as regression tests on production traffic, and the whole engagement runs under ISO/IEC 42001, the AI management system standard we are certified against.
  4. 04Can the agent run inside our own cloud account?
    Yes. We deploy into your AWS, Azure or Google Cloud account, in an India region by default for data residency, and we can use open-source models hosted in that account when data cannot leave it. You own the infrastructure, the code and the logs from day one.
  5. 05How long until a first agent is running on our data?
    A pilot on one use case runs four to six weeks and ends with a working agent, an eval report and a written go/no-go. Taking a validated pilot to production is eight to sixteen weeks depending on the integrations.
  6. 06What happens when the agent is wrong?
    Every action is traceable to the inputs, the tool call and the model version. Reversible actions are reverted; irreversible ones were never executed without an approval. The case is added to the eval set so the same failure fails a test next time.
  7. 07Do you use our data to train models?
    No. Client data is used only for the client's agent, inside the client's account. No model provider receives it for training; endpoints are chosen for that guarantee.
  8. 08How do you handle a model upgrade?
    New versions run shadow against the golden set and a sample of production traffic before they decide anything. Version pinning means an upgrade is a change request, not an accident.
  9. 09Can it run on-premises or air-gapped?
    Yes, with open-weights models (Llama, Mistral, Qwen) on your hardware or in a private subnet. Tracing, evals and gates are the same.
  10. 10Who owns the IP, prompts and evals?
    You do. Repository, prompts, golden set, infrastructure code and logs live in your accounts from the first commit.
  11. 11How does this fit the EU AI Act, FCA or RBI guidance?
    See the governance matrix on this page. We produce the logging schema, risk assessment and oversight policy each framework asks for, and we are certified against ISO/IEC 42001.
  12. 12What is MCP and why does it matter to us?
    A standard way for an agent to reach tools. It means the integrations we build are portable across models and orchestration frameworks.

Have an AI agent project in mind?

Book a 30-minute discovery call. We'll review your problem and sketch an agent architecture with the gates drawn in.

+91 912-195-7728Hyderabad, IndiaEvery brief gets a senior review. Reply within 1 business hour, 9 AM-7 PM IST.