How an external network pentest works
The real sequence of an external network penetration test: scope and rules of engagement, reconnaissance, perimeter analysis, exploitation, reporting and retest, plus what the test will not cover.
Seven phases, in the order we run them. What we need from you before day one, what happens in each phase, what lands on your desk at the end, and the things a penetration test will not tell you.
Every engagement we run moves through the same seven phases, whatever the target is. A payments API, a Kubernetes cluster and an Active Directory domain get tested with different tools, but the sequence does not change: agree the scope and the rules, map the surface, work out what a real attacker would go after, find the weaknesses, prove them, report, and come back to check the fix.
That sequence is not ours. It is the shape of the Penetration Testing Execution Standard, whose seven sections run from pre-engagement interactions to reporting, and of NIST SP 800-115, which groups the same work into Planning, Discovery, Attack and Reporting, with a feedback loop from Attack back into Discovery whenever a foothold opens new ground.
The way in has shifted. The 2026 Verizon Data Breach Investigations Report, covering incidents between 1 November 2024 and 31 October 2025, found that 31% of breaches now start with the exploitation of a software vulnerability, ahead of stolen passwords.
Nothing is sent until the scope and the rules of engagement are agreed in writing. The scope names the systems, environments and identities in play. The rules cover the rest: the testing window, the addresses our traffic comes from, how far exploitation may go, and what happens if we find evidence that somebody got there before us.
NIST ships a rules of engagement template as Appendix B of SP 800-115, and the PCI Security Standards Council’s guidance sets out the same questions. In our experience this is the part that goes wrong.
We also agree success criteria, the point at which the test is finished. PCI’s guidance is direct: defining them sets the depth of the test, and without them a tester can run past the boundaries you expected.
Reconnaissance is the first phase in which anything is sent. We build the picture an attacker builds before choosing a door: hosts and addresses, exposed services and ports, technology fingerprints, endpoints and parameters, identities and the boundaries between roles. NIST puts this first inside its discovery phase, ahead of any vulnerability analysis, and its techniques are still in daily use, from DNS interrogation to banner grabbing.
On external work this is where the surprises live: the host nobody owns, the staging copy that answers to the internet, the subdomain pointing at a service switched off two years ago. We map what we find against your own asset inventory, because the difference between those two lists is a finding in itself.
Between recon and testing we stop and do the arithmetic: which attack scenarios are realistic here, ranked by likelihood and by cost. Threat modelling is a phase of its own in PTES because it stops an engagement spending its hours on a generic checklist while the thing that would hurt goes untested.
It is also where the methodology stops being generic. The checklist comes from the standard that governs the target:
Tooling gives coverage. People give truth. Scanners run first because they are fast across a large surface; every candidate is then validated by hand. NIST puts the limit plainly: a vulnerability scanner checks only for the possible existence of a vulnerability, while the attack phase exploits it to confirm that it exists.
The findings that matter most are usually the ones no scanner has a signature for. A business logic flaw. An authorisation check applied on one endpoint and forgotten on the next. A workflow that can be replayed out of order, or completed without the step that takes the payment.
Exploitation is verification, not theatre. NIST’s attack phase exists to confirm a suspected weakness by exploiting it, and the loop back into discovery is part of the model: a successful exploit exposes new surface, analysed and tested in turn. That loop is how single findings become chains: a path traversal that only reads files stops being low severity the moment one of those files holds a credential.
Post-exploitation answers the question your board will ask: so what? We measure the reach of a foothold against confidentiality, integrity and availability, following privilege escalation and lateral movement paths as far as the rules of engagement allow, and no further.
A severity score is not a work queue. The CVSS v4.0 specification says so itself: consumers are expected to use CVSS as an input to a vulnerability management process that also weighs factors CVSS does not model. So findings are ranked on three signals, not one.
Over those three sits the signal we cannot compute for you: what the affected system is worth. A medium on the service that moves your money outranks a high on the host that serves marketing images.
The report is the product. Ours follows the structure the PCI Security Standards Council sets out, because it is legible to an assessor, an auditor and a CTO at once:
Those findings also land in our platform, where each carries its proof of concept and its business impact, and where fix verification is tracked.
Then the retest. Once you have remediated, we test the fix and either sign the finding closed or say why it is not. PCI’s guidance is precise: remediation should be completed and retested within a reasonable period after the original report, and if it stretches out, a fresh engagement is the honest answer, because the environment has moved.
Length is decided by scoping, not by a standard menu, so we will not put a number on your engagement in a blog post. What we can tell you is what moves it, because these are the questions we ask before we agree days with anyone.
Two things are worth settling before you compare one quote with another. Days quoted for testing are not calendar days: report writing, your review and the retest all sit after the last request is sent.
And depth is bounded by the success criteria the rules of engagement name, which is why PCI’s guidance treats defining them as what sets the depth of the test and keeps a tester inside the boundaries you expected. NIST says the same from the other side: a test often carries a narrow scope because of resource limitations, particularly time, while an attacker takes whatever time they need.
What a test does not cover decides whether you are buying the right thing, so it belongs in the same document as the methodology rather than in a conversation afterwards.
If any of this looks like a problem you are carrying, half an hour on a call scoping it with a senior pentester is worth more than reading another article.
Hablar con un pentester seniorPick a time that suits you. You tell us what you need and where you are, and we explain how we work and how we can help.