AI PENETRATION TESTING

Find what your AI actually exposes.

Your copilots, RAG pipelines and agents already see customer data, source code and internal policies. We prove how an attacker (or an over-trusting employee) turns that access into a breach, before it ends up in a regulator’s inbox.

How an AI pentest works
87%
of our clients renew annually
150+
organisations across multiple European countries
7/10
confirm findings their previous provider did not find
Want to know what your AI stack actually exposes? Share your main concern and your email.

Protected by reCAPTCHA. The Google Privacy Policy and Terms of Service apply.

Trusted by AI and platform teams in highly regulated environments

Identify how an attacker walks your AI to your customers’ data.

ai pentest demo
GOAL · REACH OTHER CUSTOMERS' DATA
vector store · namespace check missingauthenticated as acme · retrieves northwind docsLIVE
GET/v1/rag/retrieve?q=invoice&top_k=5
Authorization:Bearer eyJ… [tenant=acme]
doc_idtenanttitlescore
doc_8841acmeInvoice #44210.92
doc_2210globexInvoice #22100.88
doc_5512initechPayroll Q30.81
EXFILtenant_id never compared on retrieve
ACME COPILOT · CUSTOMER VIEW
Find docs related to invoice #4421
Found docs from Globex and Initech too.
Other tenants in the answer.
WHAT THE CUSTOMER SEES
AI → CROSS-TENANT PIVOT 5 HOPS · ~18 MIN
leaksystem prompt + 4 tools disclosed via polite extraction
injectPDF footnote rewrote agent goal · guardrail bypassed
rag/v1/rag/retrieve · invoices, payroll, contracts leaked
·toolrefund() auto-approved · EUR12,840 routed to attacker IBAN
·tenantsingle prompt · 18,421 users enumerated end-to-end

The AI running inside your systems is only as good as the data it is fed, and only as secure as the access around it. We test from an attacker’s point of view: what data they can manipulate to fool the model, what access it has to the AI server, what secrets sit in its container, and how it talks to the systems around it. The whole chain.

WHAT WE TEST
  • Prompt injection (direct and indirect)
  • Sensitive-data leakage via context and RAG
  • System prompt and policy extraction
  • Excessive agency in agents and tool calling
  • Vector and embedding weaknesses
  • Supply chain: models, frameworks, MCP servers
  • Data and model poisoning paths
  • Unbounded consumption and cost abuse

Why a CTO, CISO or platform lead books an AI pentest.

An AI pentest is normally booked because a release-velocity question, an enterprise customer ask, a regulator deadline, a near-miss in production or an M&A due diligence has forced it.

TAP A SITUATION
Board & customer Second opinion on incumbent provider

Someone with authority asked, and ‘we’re fine’ isn’t an answer.

When the ask is for assurance about AI exposure, you need an external, expert-led AI pentest and a deliverable that maps to business impact in language a non-technical stakeholder can read.

  • Executive summary that lands without translation
  • Findings ranked by business impact and audit exposure
  • Attestation language that satisfies enterprise procurement
  • Independence: external and expert-led throughout
Compliance & audit OWASP Top 10 for LLM Apps audit

An auditor does not accept a jailbreak demo.

With the OWASP Top 10 for LLM Apps, ISO/IEC 42001 or EU AI Act conformity work on the line, the report has to survive scrutiny: traceable scope, a methodology mapped to recognised standards and signed retest evidence.

  • Methodology mapped to the OWASP Top 10 for LLM Apps and the OWASP AI Testing Guide
  • Evidence formatted for ISO/IEC 42001, EU AI Act and regulator review
  • Execution certificate ready for the auditor folder
  • Retest evidence with timestamps, the model version pinned, and signoff
Suspected exposure / incident-driven Customer reports cross-tenant answers

Something looked off. You need to know what is still exposed.

The question is not ‘do we have findings’ but ‘are there other paths’. We focus the engagement on the suspected exposure, confirm it is closed, and surface every adjacent path an attacker could pivot to.

  • Targeted scope around the suspected exposure
  • Adjacent paths mapped: prompts, retrieval, tools, agents and the model supply chain
  • Reproducible proofs of concept for the incident-response team
  • Live findings to your responders, not the final report
Change / release New RAG source going live

Your AI ships every week. Validate the change before it does.

The cheapest moment to validate is at the change, while your team still holds the context. We test the new surface, prove which paths exist, and retest the difference after the fix to confirm the boundary held.

  • Scope tuned to what is shipping: the difference against the previous model, prompt and tool set
  • Prompts, retrieval, tool permissions and tenant boundaries pressure-tested
  • Critical findings reported live, so fixes happen mid-release
  • Retest closes the case, not the report
By AI use case Vulnerable model or framework in use

The use case decides the threat model, so the test follows the use case.

A copilot, an agent with write actions and a public chatbot each fail differently. We scope to yours and prove the path from what any user can already send to the data that should have been out of reach.

  • Threat model written for your use case, not a generic prompt-injection checklist
  • Retrieval namespacing and tenant isolation tested with real document sets
  • Tool and agent permissions tested against least privilege, write actions included
  • Models, frameworks and MCP servers reviewed as part of the surface

Ready to see what your AI actually exposes?

Book a call
Thirty minutes with an experienced AI pentester.

A seven-phase method, for your AI security pentesting.

Seven phases, in order. Each one ends with something proven, not something assumed.

Go beyond checklist AI pentesting.

Strengthen your AI posture with manual, expert-led assessments, not box-ticking jailbreak demos.

Manual, not automated.

A senior AI pentester walks every chain by hand against your models, prompts, RAG and tools. No scanner noise.

Real attacker simulation.

We prove or disprove the path from a public chat to a tool we should not be able to fire, to data we should not see.

OWASP LLM-aligned report.

Every finding mapped to the OWASP Top 10 for LLM Apps (2025): the artefact your auditor and procurement teams read.

Model-aware retest.

After a fix or a model or prompt change we retest the diff against the same version, RAG index and tool config.

Want to see what your next AI pentest with us would look like?

We at Etnia highly value our collaboration with Asperis Security.
Sergi Leno, Systems Manager · ETNIA Barcelona

Loved by engineering and security teams.

Reports your ML engineers actually read and your enterprise customers’ procurement teams actually accept. Real names, real roles, real clients.

We at Etnia highly value our collaboration with Asperis Security. Their professionalism, approachability, quick response and ability to adapt to our needs have been key in every project. The quality of service and continuous support always give us peace of mind. Without a doubt, it is a pleasure to have them as technology partners.
Sergi Leno, Systems Manager
ETNIA Barcelona
ASPERIS has worked alongside us to define and implement our cybersecurity roadmap in Microsoft 365 with a structured approach aligned to business objectives. Thanks to their advice, we took the strategic step of completing our Microsoft ecosystem and reinforcing it with CrowdStrike for advanced mobile device protection, significantly raising our security level.
Jordi Bondia, IT Director
SALVI
At NPAW we have collaborated with Asperis on various security initiatives and the experience has been very positive. We especially value their ability to adapt to our needs and the depth with which they approach each project. Results are clear, structured and useful for decision-making and continuous security improvement. We like working with Asperis for the judgment and value they bring to every collaboration. Their work has helped us strengthen our security level.
Sergi Laencina Verdaguer, CISO
NPAW
With Asperis you don’t hire a service. You hire a partner. They don’t look to bill a project. They look to establish a relationship of trust, caring about the key points that affect your organisation’s security. Professionalism, know-how and diligence.
Juan Valer Tecedor, Software Engineer
GNOSS
READY WHEN YOU ARE

Want a real AI scope?

An experienced consultant will reply.

Questions that come up before signing.

A web or API pentest treats the AI as one endpoint behind your stack. An AI pentest treats the model plus prompts plus RAG plus tools plus permissions as the product: instructions, context, retrieval, agent orchestration, model providers and the supply chain are all in scope. It is not a subset, it is a different surface. The same product can ship a clean web and API pentest and still have an indirect prompt injection in an uploaded PDF that fires a refund tool, or a RAG index that returns another tenant’s documents to a single natural-language question. Most regulated companies that ship an LLM-backed product run both: a web or API pentest tied to the front-end and gateway release train, and an AI pentest tied to model versions, prompts, tools and RAG sources.

An AI pentest covers the model plus its prompts, its context, its RAG sources, its agents and its tools, not just the endpoint your web team ships. We audit direct prompt injection, indirect prompt injection through documents your agents read, system prompt and policy extraction, RAG index integrity and cross-tenant retrieval, tool-calling permissions and excessive agency, vector store and embedding weaknesses, connector and MCP server supply-chain risks, and cost or token consumption abuse. Every finding is chained to business impact and mapped to the OWASP Top 10 for LLM Apps (2025).

An AI pentest is priced by scope, complexity and test type. We deliver a fixed-price proposal within 48 hours of our first call, with no hidden fees and no commitment to renew. A focused audit of a single chatbot with a documented RAG index sits at the bottom of the range. Multi-agent platforms, multi-tenant RAG rollouts or engagements requiring specialised harness setup for a proprietary model scale from there. The free retest is always included, against the same model version and prompt configuration.

Not without your consent. Every AI pentest ships with documented Rules of Engagement, safe-attack criteria and a critical-finding protocol: which surfaces are in scope (staging or a controlled production tenant), which tools may be fired live, and what triggers an immediate stop. On managed model providers we default to a sandboxed API key with a rate ceiling and a token budget you agree in advance, so cost-abuse testing is measured, not open-ended. Any exploitation on production requires explicit written approval and your team on standby.

Black box mirrors what a customer or a curious researcher would have: the public chat, an uploaded document and a polite question. Grey box adds your system prompts, your tool list, your RAG source inventory, a test tenant and one technical contact for questions, so we spend more of the engagement chaining and less on discovery. White box adds architecture diagrams, model configuration, prompt libraries and connector or MCP server design documents. We usually recommend grey box for the first engagement and repeat cycles as your AI product matures, and black box for a customer-audit signal or a public-facing chatbot re-baseline.

Experienced senior offensive security specialists with hands-on AI pentesting experience. Our team holds OSCP, OSCE³, OSWE, OSEP, CRTO and CRTP credentials, and our AI practice adds LLM-specific expertise (OWASP Top 10 for LLM Apps 2025, MITRE ATLAS, OWASP AI Testing Guide) plus practical familiarity with GPT, Claude, Llama and Mistral models and the common orchestration frameworks (LangChain, LlamaIndex, MCP servers). We are NASA Bug Bounty verified contributors, and our team has published research against widely deployed AI toolchains. Every engagement is scoped, executed and retested by the same specialist you spoke to on the first call. You keep talking to that same person throughout.

Ready to reduce risk across your agents?

We will review your AI stack (channels, models, prompts, RAG, agents, tools, permissions) and help you define the right AI pentest before the project begins.

Request an AI pentest proposal.
An experienced AI pentester replies within one business day. You talk to the person who runs the test.

Protected by reCAPTCHA. The Google Privacy Policy and Terms of Service apply.

Or email [email protected] directly.

OUR CLIENTS HAVE ALREADY DONE IT

We at Etnia highly value our collaboration with Asperis Security.

Sergi Leno, Systems Manager
ETNIA Barcelona

ASPERIS has worked alongside us to define and implement our cybersecurity roadmap in Microsoft 365 with a structured approach aligned to business objectives.

Jordi Bondia, IT Director
SALVI

At NPAW we have collaborated with Asperis on various security initiatives and the experience has been very positive.

Sergi Laencina Verdaguer, CISO
NPAW

With Asperis you don’t hire a service. You hire a partner.