How web reconnaissance works, and how to shrink your attack surface
By VA2PT Team, . 8 min read

Every penetration test and every real attack begins the same way: before anyone tries to get in, they build a picture of what exists. This stage is called reconnaissance. If you are learning security as a fresher, understanding recon from the defender's side is the fastest way to make yourself useful, because most organisations lose ground here long before any exploit is fired. This article explains what recon is, what a tester is actually looking for, and the concrete steps that shrink what an attacker can find.
A note on ethics and the law: the techniques below are described so you can defend your own systems and practise on machines you own or are authorised to test (for example DVWA or Metasploitable in an isolated lab, or the official PortSwigger Web Security Academy). Probing systems you do not own, without written permission, is an offence under Section 43 of India's IT Act and its equivalents elsewhere. Learn the concepts here; apply them only where you are allowed to.
The two kinds of reconnaissance
Passive reconnaissance gathers information without touching the target's systems at all. It uses public sources: search engines, certificate transparency logs, DNS records, job postings, GitHub repositories, and company filings. Because nothing is sent to the target, it is invisible and, when it uses public data, generally legal. This is also called open-source intelligence, or OSINT.
Active reconnaissance interacts with the target: connecting to ports, requesting pages, resolving hostnames against the target's own servers. It is far more revealing, but it leaves traces in logs and can trigger alerts, and it requires authorisation.
A tester moves from passive to active. So does an attacker. That order matters for defenders, because everything an attacker learns passively is information you published without meaning to.
What an attacker learns before touching you
Think of your organisation's public footprint as a map you did not know you had drawn:
- Subdomains.
app.example.com,staging.example.com,vpn.example.com. Certificate transparency logs record every TLS certificate ever issued for your domains, so a forgottenold-admin.example.comis discoverable years later. - Technology stack. Response headers, cookie names, JavaScript bundles and error pages reveal which frameworks and versions you run.
Server: Apache/2.4.49tells an attacker exactly which known vulnerabilities to try. - People. LinkedIn, conference talks and job posts name your staff and the exact technologies you are hiring for. A job ad asking for "Kubernetes 1.24 and Jenkins experience" tells an attacker what is running inside.
- Leaked secrets. Public GitHub repositories, and the commit history behind them, frequently contain API keys, internal hostnames and configuration files that were never meant to ship.
None of this requires touching your infrastructure. By the time active scanning begins, a competent attacker already has a target list.
What active scanning adds
Once a tester has authorisation, active scanning fills in the live detail: which hosts are reachable, which ports are open, and which software version answers on each. The single most common finding at this stage is not an exotic exploit; it is an outdated version. Software from 2019 answering on a port in 2026 is a finding on its own, because it means the patching process has failed somewhere.
The other common finding is things that should not be exposed at all: a database port reachable from the internet, a staging environment with production data, an admin panel with no IP restriction, a monitoring dashboard without authentication. Attackers love these because they need no skill to abuse.
Turning the attacker's checklist into a defender's checklist
Here is the useful part. Every recon technique maps directly to a defensive action.
1. Know your own subdomains
You cannot protect what you have forgotten. Keep an inventory of every subdomain and what it points to, and review certificate transparency logs for your own domains so you notice certificates you did not expect. Decommission staging and demo hosts the moment they are no longer needed, and never point them at production data.
2. Reduce what your responses reveal
Strip version numbers from server banners and error pages. A generic 500 page that leaks a stack trace, a framework name and a file path is a gift to an attacker. In most web servers this is a small configuration change:
# nginx: hide the version in Server headers and error pages
server_tokens off;
3. Close the ports that should not be open
The most reliable security control is not a clever tool; it is a small attack surface. Databases, caches, message queues and admin interfaces should never be reachable from the public internet. Put them in private subnets, reachable only through a bastion or a VPN. A quarterly review of "what is listening, and does it need to be?" prevents more incidents than any scanner finds.
4. Keep secrets out of code
Assume every public repository, and its full history, will be read by someone hostile. Use a secrets manager, add pre-commit scanning (tools such as gitleaks) to catch keys before they are committed, and rotate any credential that has ever touched a repository.
5. Patch on a schedule you can prove
Outdated software is the finding that appears in almost every VAPT report. A patch process with owners and deadlines, tracked like any other work, closes off a huge share of what recon turns up.
How this maps to a real engagement
When VA2PT runs a vulnerability assessment and penetration test, reconnaissance is the first phase, and the report almost always opens with attack-surface findings: an exposed staging server, a leaked key, an unpatched service, a subdomain nobody remembered. These are rarely glamorous, and they are almost always the cheapest and highest-impact things to fix. A team that has already done the five steps above starts an engagement from a much stronger position, and the test can spend its time on the deeper issues instead.
What to do this week
- Build the list of every domain and subdomain you own, and what each one serves.
- Check your web server and application error pages for leaked versions and stack traces.
- Review your cloud security groups and firewall rules for any database or admin port open to
0.0.0.0/0. - Add secret scanning to your repositories, and rotate anything it finds.
- Write down your patching process, with an owner and a cadence, even if it is one page.
Reconnaissance is not magic. It is the disciplined collection of things you left visible. Learn to see your organisation the way an attacker does, and most of the work of defence becomes obvious.
- ethical-hacking
- freshers
- reconnaissance
- attack-surface
- blue-team
- vapt