Legal entity and trading history
Who you are contracting with, and for how long they have traded.
A practical, evidence-first checklist for deciding whether an AI consultancy is ready to touch production systems, sensitive data or consequential business workflows.
Evidence over claims.
Judge an AI consultant on verifiable production evidence, security boundaries, operational responsibility, measurable outcomes and accountability, not on demos, decks, certifications or hype.
No consultancy will answer every question perfectly, and not every question carries the same weight in every engagement. Weak or unverifiable answers should change your risk assessment, particularly when the consultant will have access to production systems, sensitive information or consequential workflows.
A prototype, a pitch or an internal demo is not a deployed system. The useful answer names systems that ran against real workloads, and says who carried them once they did.
Ask for: where it was deployed; how long it has operated; what the consultant was responsible for; how it was supported.
Launch is the cheapest part of a system’s life. What matters is whether it is still running, still supported, and whether the consultant stayed accountable after go-live.
Ask for: time in production; current support arrangement; who is accountable now.
A measured outcome and a projected ROI are different claims. Ask which one you are being shown, how it was measured, and against what baseline.
Ask for: revenue, cost, time saved or throughput; error or risk reduction; the baseline and the measurement method; whether the figure is measured or projected.
Systems in production fail. A consultancy that can describe its incidents, wrong assumptions and integration problems, and what it changed as a result, is showing operational maturity. A spotless history is not by itself evidence of one.
Ask for: a specific incident or failure; its cause; what changed in the system or the process afterwards.
The handover is where responsibility most often goes missing. Agree before signing who operates the system, who responds when it fails, and what support the engagement includes.
Ask for: the operator; the incident responder; included support and its term; the split of responsibilities between consultant and client.
Every model, API and provider will at some point be unavailable, degraded or changed. The question is whether the design already decided what happens then.
Ask for: retry and fallback behaviour; whether it fails open or fails closed; degraded and manual operating modes; how provider outages and model changes are handled.
Guidance, policy, authorisation, enforcement, monitoring and logging are six different things. A policy document describes what should happen; an enforcement point is the place in the running system that stops what should not. Ask where the system prevents an unauthorised action, and whether anything can reach the action without passing through that point.
Ask for: the enforcement point, named in the architecture; what happens when it denies; whether it can be bypassed.
In practice: One mandate. One authorised action.
An agent’s risk is bounded by what it can reach. Ask for the inventory in writing, before the system is connected to anything.
Ask for: data, files and databases; APIs, tools and production systems; payment, email and messaging systems; shell access and external network destinations.
Shared and long-lived credentials make an incident hard to contain and hard to attribute. Each agent or integration should have an identity of its own, scoped to what it needs.
Ask for: API keys and service accounts; OAuth scopes and token lifetimes; rotation and secrets management; whether credentials are shared or per agent.
An approval step only controls something if it cannot be skipped, and if the action that runs afterwards is exactly the action that was approved.
Ask for: which actions trigger approval; who may approve; whether approval can be bypassed; whether execution resumes under the exact approved authority; whether both approval and execution are auditable.
After an incident you need the chain: human, agent, policy or authority, tool, resource, result. If any link is missing from the record, the question cannot be answered later.
Ask for: what is logged at each link; how long records are retained; whether records can be altered; whether they can be verified independently.
In practice: Agent Under Oath v0.1
Untrusted content reaches agents through email, web pages, documents and tool output. Testing reduces these risks and shows where they remain; it does not eliminate them, and a credible answer says so.
Ask for: direct and indirect prompt-injection tests; tool-output injection; data-leakage and excessive-permission tests; privilege escalation and compromised tool or MCP integrations.
Revocation is the control used under the most pressure. Ask for it as a procedure with a measured time, not as a capability.
Ask for: where the kill switch is; who can use it; which credentials must be revoked; how long revocation takes; whether existing sessions or tokens survive it.
Your data passes through the consultant, its subprocessors and its model providers. Each one is a place to ask about retention, location and use.
Ask for: collection, retention and storage location; subprocessors and model providers; whether data is used for training; deletion, backups and access controls.
Every material claim above can be supported by something that exists outside the conversation. Where it cannot be, the claim is an assurance, and should be weighed as one.
Ask for: architecture diagrams; deployment evidence and operating metrics; test results and security controls; incident and remediation records; contractual responsibilities and audit evidence; references, where available.
Standard business due diligence is a useful risk signal. It does not substitute for the technical evidence above: a well-insured, long-established firm can still ship an agent with no enforcement point.
Who you are contracting with, and for how long they have traded.
Production and customer work comparable to what you are buying.
Professional indemnity and cyber cover, where the engagement warrants it.
Liability, intellectual property, termination, and the terms under which your data is processed.
Litigation or material disputes, where disclosure is appropriate.
Engagements that ended early, where relevant.
Clients you can speak to, where they are available.
Who will do the work, and whether they have the time to do it.
Only where the engagement actually requires them.
ThePlus Tech aims to answer these questions in its own engagements, and expects buyers to ask them. It does not claim to satisfy all fifteen today: the evidence on this site is from repository CI, not from customers, and the gaps are discussed during scoping rather than left for the buyer to find.
Production, pilot, demonstration, experimental and research work are named as what they are. Every product screen on this site carries the environment marker it was captured in, and every case study lists its limits beside its measurements.
Tests, logs, control evidence, architecture and measured outcomes are preferred to statements. The figures on this site are copied from the product repositories’ own trackers and test runs, not from customers, and the site says so.
Finding a risk does not mean it has been controlled. A finding is reported as a finding until the control that closes it exists and has been tested.
Policy documentation alone is not authorisation. A claim that an action is controlled names the place in the running system that stops it.
If something was not assessed or not verified, it is said, at the same weight as what was.
Where this applies to a specific engagement, see the Security Questionnaire Rescue Sprint and the case studies for TrustLedger, CyberGuardPlus and the Agent Clearing Network.
Tick these off during the evaluation. The page prints cleanly from the browser.