Concept and method · Blog
AWS Pentest - What can an identity actually reach in AWS?
The name of a role in AWS does not show what that role can actually do: the effective decision comes out of the intersection of identity policy, permissions boundary, session policy, service control policy, resource and trust policy, and conditions. This playbook sets out end to end a penetration testing method that moves from reading declared permissions to measuring effective decisions: safe test levels, the evidence model, a sixteen-phase testing flow, cross-service attack chains, risk rating and rollback.
The most common mistake in cloud security testing is confusing checking a setting visible in the console with verifying what result that setting produces. A role may be named “read-only”; but once the trust policy, resource policies and conditions are evaluated together, what can it actually read? A bucket may look “private”; but do the account-level public access block, the bucket policy, the ACL and the access point policies all protect it at once, or only one of them?
This playbook has a single guiding question, and it determines every test: not “is the permission declared?” but “what business outcome can this identity actually produce once all identity and resource policies are evaluated together?” In AWS, real risk does not come from a single permission; it comes from the combination of identity policies, resource policies, organization guardrails, service roles, network paths, encryption keys and event-driven automation. Two permissions that are harmless on their own can, used together, turn into cross-account access or code execution under a highly privileged service role.
Below you will find, for every link in the path from a low-privilege principal’s initial access to cluster, account or organization impact, which safe method to test it with. This article is only for environments you own or have written authorization for; the controls do not depend on a particular tool but on AWS API behavior, policy evaluation logic and observable outcomes.
1. The purpose of the test: which questions must be answered with evidence?
These questions drive your test from beginning to end, and each must be answered with evidence rather than assumption:
- Are the organization, account, region and service boundaries fully visible?
- How are human, workload and automation identities authenticated?
- Which resources can a low-privilege principal actually reach?
- Does one permission combine with permissions in other services to produce privilege escalation?
- Are cross-account trust relationships limited to the expected principals and conditions?
- Are internet-facing resources genuinely necessary, and are their protective layers working?
- Does an encryption key’s policy grant broader access than the data it encrypts?
- Are security logs separated from any account an attacker could reach, and are they immutable?
- Is a risky action detected, turned into a meaningful alarm, and delivered to its owner?
- Can the resources, external dependencies and cost effects created during the test be fully reversed?
2. Method first: the safe testing model
“Touching production by accident” carries an extra dimension in AWS: cost. So clarify how you will test before you decide what to test.
2.1 Test levels (S0–S4)
| Level | Approach | Example activity | Risk |
|---|---|---|---|
| S0 | Document and architecture review | Landing zone, data flow, guardrail design | None |
| S1 | Read-only discovery | Inventory, policy resolution, exposure and log review | Very low |
| S2 | Isolated verification | Separate test account, canary data, time-boxed role, alarm generation | Low |
| S3 | Controlled advanced verification | Cross-service permission chain, temporary resource policy, event triggering | Medium to high |
| S4 | Resilience / disruption testing | Quota, load, restore from backup, component loss | High |
S3 and S4 are not part of the default scope. Downloading real data, decrypting a production secret, creating long-lived access keys, stopping logging, deleting backups, changing production code, creating resources with uncontrolled cost, or removing a guardrail are not done without separate and explicit authorization.
2.2 The “stop at the first safe evidence” rule
You do not have to exploit a risk to show its impact; minimum evidence is enough:
- For data access: a unique canary object rather than a real record.
- For secret access: a synthetic secret and a one-way digest rather than the real value.
- For decryption: test data generated with a test key rather than production ciphertext.
- For service role passing: a synthetic role that reaches only a canary resource, rather than an administrative role.
- For an event chain: a test function producing fixed, harmless output rather than a real function.
- For cross-account access: two dedicated test accounts rather than a production account.
The chain stops once canary evidence is obtained. Being able to reach data is not a reason to go on to the real data.
2.3 Stop conditions
Active testing stops and the relevant parties are notified if any of the following appears: unexpected growth in billing or resource consumption; error, latency or capacity impact on production traffic; movement into an out-of-scope account, region or resource; contact with real personal data, credentials or a production secret; a test identity turning out to hold broader permissions than expected; security automation starting isolation, deletion or access revocation; a rollback step failing; an unexpected change in central logging or a security control.
2.4 Test identity and test account standard
Carrying out the whole assessment with a single administrative identity makes it impossible to measure effective trust boundaries, because you are trying every door with a key that already opens it. Test at least these profiles separately: a read-only auditor with organization visibility; a low-privilege developer; a deployment automation role; an application runtime role; an account administrator; a security and log account reader; an emergency access role; a synthetic test role standing in for an external account. Identities must be used through federation, protected by MFA, with short sessions and unique session names.
Where possible, S2 and above are performed in a dedicated test account inside the production organization. That account: holds no production data; is connected to central logging and detection; inherits the same guardrails as production; has strict budgets and quotas; keeps external connectivity limited; uses tags and lifecycle rules that expire test resources automatically; and can be shut down in full when needed.
3. The evidence model: how do you record a finding?
Keep this record for every check:
Test ID:
Time (UTC):
Test level (S0–S4):
Source account / region class:
Test principal:
Service and resource type involved:
Policy examined or request sent:
Expected safe outcome:
Observed outcome:
Verdict: PASS / FAIL / PARTIAL / BLOCKED / N/A
Masked evidence or evidence hash:
Cost and service impact:
Rollback step and its verification:
These are always masked: account numbers and real ARNs; user and role names; IP addresses, domain names, bucket names; access keys, session tokens, signed requests; secret and parameter values; database endpoints and snapshot names; personal or business data in logs.
Verdicts: PASS (safe behavior demonstrated), FAIL (risk or precondition safely confirmed), PARTIAL (some policy layers protect and others do not, or only an alarm is produced), BLOCKED (could not be completed because of missing permission, visibility, dependency or approval), N/A (not applicable given the architecture).
A critical nuance: AccessDenied is not always PASS. A denial may mean the control is sound; but if the test identity is not even permitted to see the inventory, the verdict is really BLOCKED. Do not count a control as passed before confirming that the denial came from the right policy layer and for the expected reason.
4. The end-to-end testing flow
Phase 0 Organization and trust boundary map
The map comes before the commands. The goal is not a service list but a graph of “which principal can pass through which policy layers to reach which business data or runtime role?” Correlate onto it: the organization/OU/account hierarchy; management, security, log archive, network and workload accounts; organization policies and delegated administrator roles; human and workload identity sources; cross-account trust relationships; central network, DNS and internet egress paths; data classes, KMS keys and backup targets; the central audit/config/threat detection pipeline; the CI/CD and IaC path.
Phase 1 Global and multi-region discovery (S1)
| Test ID | What is tested | Finding criterion |
|---|---|---|
| AWS-DISC-01 | Account and region inventory - do not settle for the default region, examine every active region | An active region left out of scope, an unowned account, an unclear data boundary |
| AWS-DISC-02 | Resource inventory and tag governance (owner/env/data-class/criticality/expiry coverage) | A resource whose owner or data class is unknown; a control driven by tags being trivially bypassed |
| AWS-DISC-03 | External exposure - public IPs, internet-facing load balancers, public APIs, CDN origins, DB endpoints, management endpoints; match DNS to real resources | An unowned external endpoint, a dangling alias, an unverified public service |
| AWS-DISC-04 | Principal and credential inventory (owner, last use, creation, expected use) | An unowned key, a shared identity, a long-lived credential |
Note: an untagged resource is not automatically counted as “low risk”.
Phase 2 Organization guardrails and the account baseline
AWS-ORG-01 Root user. MFA, absence of access keys, non-use in daily work, current contact details, a usage alarm. Opening an active root session is not a default test; configuration and audit evidence is sufficient.
AWS-ORG-02 SCP coverage. Compute the effective guardrails of every OU: region restriction, log protection, security services that cannot be turned off, public access and root limits. Finding: a critical deny guardrail applied to only some accounts, or a broad exception.
AWS-ORG-03 Delegated admin and the management account. Delegated administration for security services, human access in the management account, workload automation being unable to reach organization management, and account creation/move permissions.
AWS-ORG-04 Default security baseline. The controls a new account inherits automatically: central audit, config recording, threat detection, public access guardrails, default encryption, budget and contacts, management of default networks. A landing zone control must be verified not only on existing accounts but across the lifecycle of a synthetic new account (S3 plus cost approval).
Phase 3 Authentication, IAM and privilege escalation
This is the heart of AWS testing, and the one golden rule here is: IAM is not a list of identity policies; the effective decision is the intersection of many layers.
Identity policy
∩ permissions boundary
∩ session policy
∩ service control policy (SCP)
∩ resource policy / trust policy
∩ conditions
− explicit deny
= effective permission and business impact
You cannot say what a principal “can do” without solving this formula. All the tests below probe that intersection from different angles.
AWS-IAM-01 Federation, MFA, sessions. Human access coming from a central identity, phishing-resistant MFA for privileged roles, session duration and re-authentication, quick removal of a disabled user, approval of permission set assignments.
AWS-IAM-02 Local users and access keys. Justify local users carrying a console password or keys. Age alone is not enough; last use, service, source IP and owner are assessed together. Two active keys, a never-used key and an old key left enabled after rotation are all flagged. Finding: a long-lived key on a human user, a shared user, an unowned credential in automation.
AWS-IAM-03 Trust policy analysis. Wildcard principals; overly broad trust to a root principal; ExternalId, organization or resource conditions for external accounts; audience/subject limits for web identity; SourceArn/SourceAccount for service principals; session tag behavior; role chaining and maximum session duration. Remember: the question “who can assume this role?” is answered not by the trust policy alone but together with the assume permission on the other side.
AWS-IAM-04 Wildcard and unconditional permissions. Action:*, Resource:*, NotAction/NotResource are examined in context. The decrypt, secret and log permissions of a role named “read-only” are assessed separately. Whether tag-based conditions rely on tags the principal can modify is checked.
AWS-IAM-05 Direct policy manipulation. These permission families count as separate attack paths: writing an inline policy; attaching a managed policy; creating a new policy version and making it default; changing a trust policy; removing or changing a permissions boundary; changing a group or permission set assignment; creating a new access key or login profile. At S1, effective permission analysis only; at S2, a dedicated test principal and an inert canary policy - real escalation is never performed.
AWS-IAM-06 Role passing combined with service creation. A principal’s ability to pass a role to a service (PassRole) is not assessed on its own; its intersection with these is computed: creating or updating a function; creating an instance, task, job, notebook or build; creating a stack or deployment; automation documents or remote commands; event targets, schedulers, pipelines; container task definitions. Finding: a low-privilege principal being able to pass a role stronger than its own to a service that runs code it controls.
AWS-IAM-07 Implicit privilege escalation. Permissions that contain no administrative action yet carry high risk:
| Permission type | Combined risk |
|---|---|
| Updating function code/configuration | Acting with the function execution role’s permissions |
| Changing instance user data / launch template | Moving to the instance profile’s permissions |
| Running automation/commands | Commands on instance identities |
| Changing a stack/template | Creating resources with the deployment role |
| Changing an event rule/target | Triggering a target that holds a stronger role |
| Changing a data pipeline/build definition | Running code with the build/job role |
| Writing a resource policy | Granting access to oneself or an external principal |
| Creating a KMS grant | Enabling access to encrypted data |
| Changing a secret policy | Opening the secret to another principal |
AWS-IAM-08 Cross-account access matrix. Record every trust relationship with these columns, and do not automatically treat in-organization access as safe:
Source account class → Target account class → Principal type → Role → Conditions
→ Session duration → Effective permission → Data reached → Alarm state
Examine the development→production, workload→log archive and vendor→management role paths in particular.
AWS-IAM-09 Unused and shadow permissions. Long-unused roles and policies, roles never assumed, bindings belonging to former staff, multiple paths granting the same administrative effect, and use of the emergency role in normal work.
Phase 4 Audit, configuration and detection
Test the ability to see as thoroughly as you test prevention, and confirm that security logs live in a trust domain the attacker cannot reach.
AWS-LOG-01 Organization-wide, multi-region audit. An organization trail covering every account and region, global service events, the necessary data event coverage, log file validation, delivery to a separate log archive account, and an alarm on delivery failure or delay.
AWS-LOG-02 Log store protection. The service that writes to the log bucket is separated from the analyst role that reads it; workload accounts must not be able to delete logs, change the lifecycle or disable the key. Versioning, the need for Object Lock, encryption, a narrow resource policy.
AWS-LOG-03 Critical control changes. Confirm that these operations are both recorded and turned into an alarm: stopping or deleting a trail, changing the config recorder, disabling threat detection, changing a security service administrator, changing an SCP or log bucket policy, KMS disable or deletion schedule, relaxing public access, root or break-glass use.
AWS-LOG-04 Configuration drift. Configuration history for critical resource types, new accounts and regions being brought into coverage automatically, tickets or automatic remediation for non-compliant resources, and the business impact and rollback of that remediation.
AWS-LOG-05 Detection verification. In the S2 account, generate harmless events with a unique session: a failed role assumption, a denied canary secret access, an attempt to add a public policy, use of a disallowed region, adding a broad rule to a test security group and reverting it immediately, a limited anomalous API pattern for a synthetic key. The API event, central log, detection, alarm, ticket and time to reach the owner are each measured separately.
Phase 5 VPC, network and edge security
On the network, no single rule decides anything: source IP + route + NACL + destination listener + host firewall are assessed together. A broad security group is not by itself proof of internet reachability, though it may be a gap in defense in depth.
- AWS-NET-01 VPC/subnet architecture: public/private/isolated/management separation, route tables and IGW/NAT paths, shared VPCs, transit gateways, peering, VPN, overlapping CIDRs, prod↔dev crossing paths.
- AWS-NET-02 Security group effective access graph:
0.0.0.0/0and::/0rules, management ports, SG reference chains, broad egress, unowned groups, the LB→application→DB path. - AWS-NET-03 NACL and stateless behavior: inbound and outbound together, ephemeral port ranges, IPv4/IPv6 equivalence, bypasses caused by wrong rule numbers.
- AWS-NET-04 VPC endpoints and endpoint policy: private workloads reaching a public service endpoint, and the endpoint policy permitting only the necessary principals and resources. The existence of an endpoint does not mean a data perimeter is in place; identity, resource and endpoint policies are tested together.
- AWS-NET-05 Flow log coverage: coverage of critical VPCs, subnets and interfaces, recording of both accept and reject, a central destination, queryability, and permissions to stop or delete flow logs.
- AWS-NET-06 Load balancer and TLS: scheme selection, listener rules, HTTP→HTTPS redirection, modern TLS policy and certificate lifecycle, the need for backend re-encryption, access logs, information leakage from the health check endpoint.
- AWS-NET-07 WAF/CDN/origin: WAF association and coverage, rule exceptions, rate-based controls, blocking direct access to the origin behind a CDN, host header and cache behavior. An active application attack is not in default scope; what is measured is the placement of the cloud-side control chain and the result it returns for safe canary requests.
- AWS-NET-08 DNS and certificates: public/private zone separation, dangling aliases and abandoned records, zone delegation permissions, the need for DNSSEC, certificate validation records and automatic renewal, use of wildcard certificates.
Phase 6 Compute and instance security
- AWS-EC2-01 Public exposure: public IPv4/IPv6, public subnet routes, SG/NACL paths, direct access to an instance expected to sit behind a load balancer, source networks for management ports.
- AWS-EC2-02 Instance Metadata Service: token requirement (IMDSv2), hop limit, whether the metadata endpoint is genuinely needed, access in container and reverse proxy scenarios. In active SSRF verification no real application or credential is used; a synthetic instance profile and a test endpoint reaching only a canary resource are used, and temporary credentials are never recorded as evidence.
- AWS-EC2-03 Instance profile and role scope: the justification for each instance’s role need, the same role being shared across different trust levels, secret/decrypt/object/remote-management permissions, role assumption and cross-account access.
- AWS-EC2-04 User data, launch templates, images: no secrets or tokens in user data, permissions to change launch templates, image ownership, sharing, currency and provenance, and a new instance inheriting the security baseline automatically.
- AWS-EC2-05 Disks and snapshots: default EBS encryption, volume and snapshot keys, public or cross-account snapshot sharing, copy and export paths, volume persistence after an instance is deleted. A real snapshot is never restored; this is done at S2 with a synthetic disk and a canary file.
- AWS-EC2-06 Remote management and patching: central, recorded sessions instead of direct SSH/RDP, session logs going to a separate account, patch baselines and non-compliant or unmanaged instances.
- AWS-EC2-07 Lifecycle: termination/stop protection, auto scaling health replacement, cleanup of unused elastic IPs, volumes, snapshots and images.
Phase 7 Object storage (S3) security
Public access is not a single setting but a layered decision. Examine these layers together: the account-level public access block, the bucket-level public access block, the bucket policy, bucket and object ACLs, access point and multi-region access point policies, website hosting, organization and VPC endpoint conditions. A public access analyzer result is not sufficient on its own; an anonymous canary object request is made only against an explicitly designated test bucket and with S2 approval.
- AWS-S3-02 Bucket policy: wildcard or cross-account principals, the secure transport requirement, organization / source VPC / source ARN / principal tag conditions, deny guardrails, wrong or controllable context keys.
- AWS-S3-03 Encryption and key impact: default SSE, data classes requiring a customer-managed key, encryption enforcement on upload, encryption of replication and inventory output, and a KMS key policy creating access beyond the bucket policy.
- AWS-S3-04 Versioning, Object Lock, deletion resistance: delete marker behavior, the need for Object Lock, lifecycle rules deleting older versions, replication and backup, and the same principal being able to delete the data and its protection together (the ransomware scenario).
- AWS-S3-05 Critical bucket classes: in audit log, IaC state, CI/CD artifact, backup and sensitive data buckets, read, write, policy-lifecycle and key management must be separated across different principals.
- AWS-S3-06 Event notification attack surface: changing the notification configuration, the resource policy of the target function/queue/topic,
SourceArn/SourceAccountvalidation, event filter coverage, the target execution role’s permissions. At S2, a test bucket, a test target producing fixed output and a canary role; production notifications are never changed. - AWS-S3-07 CORS, presigned URLs, access logs: broad origin/method combinations, presigned URL duration and permission source, separation of write and read on the access log destination, object ownership, sensitive object names appearing in logs.
Phase 8 KMS and cryptographic boundaries
A decrypt permission is not a finding on its own; the real question is whether that permission is connected to a data path:
Principal → decrypt/grant permission → key
→ resources encrypted with that key
→ permission to read the ciphertext or the resource
→ real data impact
The risk of a principal that cannot reach the ciphertext is not the same as that of one which can also read the secret, object, snapshot or database data.
- AWS-KMS-01 Key policy and effective principals: the key policy, IAM policies and grants are resolved together; how much authority the use of the account root principal delegates to IAM is understood; wildcard and cross-account use, and the administrator/user separation.
- AWS-KMS-03 Conditions and the data perimeter:
ViaService, encryption context, caller account/organization/alias conditions; the need for direct decrypt; multi-region keys and replica policy consistency. - AWS-KMS-04 Grant management: grant create/retire/revoke permissions, use of grant constraints, service-originated temporary grants, unused or broad grants.
- AWS-KMS-05 Lifecycle and separation of duties: rotation, disable and deletion schedule permissions, separation of the key administrator from the user who decrypts data, protection of backup and log keys from workload accounts. Active decrypt testing is never done with a production key or ciphertext; a dedicated test key and canary ciphertext are used.
Phase 9 Secret and configuration data
- AWS-SEC-01/02 Inventory and resource policy: secret owner, consuming workload, rotation plan, replication and KMS relationship; wildcard or unexpected principals, organization and source conditions, the intersection of resource and identity policy, an external account also being able to use the key.
- AWS-SEC-03 The real behavior of rotation: do not settle for a “rotation enabled” label; the last successful rotation, failed attempts, revocation of the old credential and the application using the new value are all verified.
- AWS-SEC-04 Secrets kept in the wrong place: function environment variables, instance user data, image layers, build logs, IaC state, object metadata and tags, plain-text configuration. Values found are never copied into evidence; only the location, type and rotation need are recorded.
- AWS-SEC-05 Canary secret verification: a synthetic secret must be readable only by the designated test role; a low-privilege human role, another workload role and an external account must all be denied. The alarm chain is measured for both successful and denied access.
Phase 10 Managed databases and data services
PubliclyAccessible: false is not proof of private access on its own; peering, transit and shared network paths are taken into account.
- AWS-DB-01 Public reachability: publicly accessible plus subnet route plus SG/NACL plus DNS plus proxy path, together.
- AWS-DB-02/04 Encryption and TLS: instance/storage/log/replica/snapshot/export encryption, KMS access, client TLS enforcement, certificate validation, security parameters, the default parameter group.
- AWS-DB-03 Authentication: use of a static admin password, IAM or database authentication, secret rotation, master user limits, application role privileges.
- AWS-DB-05 Snapshot and export sharing: public snapshots, cross-account sharing, export destination and role, the data class reachable through a restore. A production snapshot is never restored; a test database with synthetic data is used.
- AWS-DB-06/07 Backup and NoSQL/cache: retention, PITR, deletion protection, multi-AZ, cross-region and cross-account backup, restore drills; for NoSQL and cache, resource policies, stream consumers, encryption and TLS, public or peered paths, cache authentication.
Phase 11 Serverless and event-driven architecture
- AWS-SRV-01 Execution role: a separate, least-privilege role for every function; the same role not being shared across different data classes; secret, decrypt, network and cross-account permissions; the trust policy.
- AWS-SRV-02 Invoke and resource policy: public function URLs and API integrations, wildcard principals,
SourceArn/SourceAccount, cross-account invoke, access to async destinations. - AWS-SRV-03 Code/config change permissions: the combination of code/layer/env/handler/runtime update permissions with the execution role, code signing and artifact provenance, the deployment principal’s boundary.
- AWS-SRV-04/05/06 Environment, concurrency, VPC: secrets in env, ownership of shared layers, data residue in ephemeral storage; reserved concurrency, retry/DLQ, recursive invocation protection, timeout and memory, and the cost and availability impact of unbounded events; whether the function needs a VPC and what its egress path is.
- AWS-SRV-07 Event source trust: the identity of the sources, event filter coverage, replay and duplicate behavior, the poison message and DLQ path, the target role and resource policy.
Phase 12 Messaging and integration
- AWS-MSG-01/02 Queue and topic policies: public or cross-account send-receive-publish-subscribe,
SourceArn/SourceAccount, TLS enforcement, SSE, unapproved external subscriptions, DLQ access, filter behavior. - AWS-MSG-03 Event bus and rules: cross-account
PutEvents, event bus resource policy, permission to change rules and targets, target role passing, archive and replay permissions, secrets in event content. - AWS-MSG-04 Event injection verification: at S2, a synthetic event goes only to a test bus, queue or topic; the target produces a fixed canary output; production consumers, data and target roles are never used.
Phase 13 Container and artifact services
This section focuses on the AWS control plane; in-cluster detail is handled by a separate Kubernetes playbook.
- AWS-CON-01 Registry: repository policy and cross-account pull/push, tag immutability, scan on push, lifecycle, encryption and replication, use of public repositories.
- AWS-CON-02 Role separation: the control/agent execution role is separated from the application task role; secret and registry access only for the role that needs it; the combined effect of task definition update plus role passing.
- AWS-CON-03 Runtime and network: public IPs, network mode and SGs, privileged mode and capabilities, host volumes and sockets, read-only root filesystem, log and secret injection.
- AWS-CON-04 Cluster control plane: endpoint access, control plane logs, workload identity trust conditions, node role scope, cluster and node group permissions.
Phase 14 Backup, disaster recovery and deletion resistance
- AWS-BCK-01/02 Plans and vaults: critical resources being included in the plan, tag and selection errors, retention and lifecycle, cross-region and cross-account copies; vault policy, the need for Vault Lock, recovery point deletion permissions, separation of the backup operator from the workload administrator, and whether deleting a key also affects the backups.
- AWS-BCK-03 Restore testing: a restore is not counted as successful on “job completed” alone. On an isolated network, with synthetic data, data integrity, usability by the application, the required KMS access, DNS and dependency behavior, and RTO/RPO are all verified.
- AWS-BCK-04 Ransomware scenario: whether the same principal can delete production data, snapshots, recovery points, logs and KMS keys together is analyzed on the permission graph. No actual deletion is performed.
Phase 15 Supply chain and infrastructure as code
- AWS-SC-01 CI/CD identity: federated workload identity instead of long-lived keys, repository/branch/environment conditions, separation of the pipeline role from the deployment role, cross-account deployment trust, traceability through session tags.
- AWS-SC-02 Artifact integrity: provenance of the build output, signature and digest verification, artifact write permissions, the same principal not performing build, approval and production deployment, use of mutable artifacts.
- AWS-SC-03/04 State and stack roles: state encryption, locking and versioning, the principals that read and write state, secrets in plan files; the stack execution role, the combined effect of template and parameter update permissions, custom resource and macro trust, deletion protection and rollback behavior.
Phase 16 Cost and resource abuse
- AWS-COST-01 Budgets and anomalies: budgets per account, project and service, cost anomaly alarms, the alarm owner being current, and response time to use of a new region or an expensive service.
- AWS-COST-02 Creation limits: permissions that can create large instances, GPUs, high concurrency or data transfer cost, service quotas and guardrails, automatic expiry rules. Active cost testing is S4 and is not performed without a low budget, a short duration and automatic shutdown.
5. Cross-service attack path scenarios
Once you have learned to test individual findings, the real skill is chaining them. The goal is not to take over the account but to show the combined effect of permissions with minimum, reversible evidence.
Scenario 1 Role passing → service → a stronger runtime identity. The roles a low-privilege principal can pass and the code-running service resources it can create are computed; at S2, a synthetic role that can read only a canary object is used; the workload confirms canary access and stops; the resources are removed. Root cause: the PassRole permission was never constrained by role ARN and target service conditions, and it combined with a resource creation permission.
Scenario 2 Object event → function → execution role. The intersection of notification change and object write permission is derived; the source conditions of the target function’s resource policy are examined; at S2, a test bucket, function and canary role; a synthetic event produces a fixed canary log; production is not modified.
Scenario 3 KMS decrypt → encrypted data. The keys a principal can use are matched against the resources encrypted with them that the principal can also read; canary ciphertext generated with a test key is used instead of production data; once decryption succeeds, the data path is proven.
Scenario 4 Instance metadata → temporary role → lateral access. Token and hop-limit settings are reviewed read-only; the resources the instance profile can reach are derived; at S2, a synthetic instance, canary role and canary object; temporary credentials are never recorded, only canary access is confirmed.
Scenario 5 Function update → execution role. The principals that can update function code or configuration and the effect of the execution role are determined; production functions are never modified; the same permission pattern is verified at S2 with a synthetic function and canary role.
Scenario 6 Stack/template control → deployment role. Stack creation/update and execution role passing are analyzed together; a synthetic stack creates only low-impact canary resources that expire automatically; rollback and the deletion of all dependent resources are verified.
Scenario 7 Resource policy → cross-account data access. The principals that can write resource policies and the guardrails blocking external access are examined; a canary object or secret is used between two dedicated test accounts; the policy and resource are removed immediately after the evidence.
Scenario 8 Snapshot sharing → data copying. Snapshot attribute changes and KMS access are computed together; production snapshots are never used; a test snapshot with synthetic data is shared to a dedicated account, canary integrity is confirmed and the copies are deleted.
Scenario 9 Bypassing log protection. The workload administrator’s trail, destination, log bucket, lifecycle and KMS permissions are derived; production logging is never stopped; a change expected to be blocked is attempted on a dedicated test trail; the explicit deny, the audit event and the alarm are recorded together.
6. Risk rating
Assign severity not by the name of the API action but by weighing these dimensions together:
| Dimension | Question to ask |
|---|---|
| Initial access | Anonymous, internet, low-privilege user, workload or administrator? |
| Policy layers | How many controls have to fail at the same time? |
| Blast radius | A single resource, a service, an account, the organization, or a connected system? |
| Data impact | What is the outcome for confidentiality, integrity and availability? |
| Chainability | Does it combine with role passing, decryption, a resource policy or an event? |
| Persistence | Can it be made durable through new credentials, trust or a pipeline? |
| Detectability | Is there an audit record and an alarm; can the attacker affect them? |
| Cost impact | Can resource consumption or data transfer cause damage? |
The same technical finding is rated differently by context: PassRole is not critical on its own, and becomes critical when combined with a strong role and a service that runs the attacker’s code. Decrypt is limited without access to ciphertext and rises when combined with reading a secret, object or snapshot. A public IP is not verified exposure by itself; it is assessed together with the route, security group, listener and application authentication. A wildcard on an administrative role may be a design decision; the same wildcard granted to a low-privilege role is an entirely separate finding.
7. Rollback and closure
The test ends when the environment is left as clean as it started. In AWS that requires sweeping several regions and hidden dependencies:
- Stop active test requests and event generation.
- Remove scheduler, rule, target, subscription and notification links.
- Terminate functions, tasks, instances, stacks and other compute resources.
- Revert temporary role, policy, trust, grant and resource policy changes.
- Delete canary secrets, objects, keys, queues, log groups and database resources.
- Check snapshots, images, backups, replicas and cross-account copies.
- Remove security group, endpoint, DNS, certificate and load balancer dependencies.
- Revoke test access keys, sessions, certificates and federated assignments.
- Search again for leftover test resources in every region, using the tag inventory.
- Compare billing, audit, config and detection records against the starting baseline.
A stack appearing deleted does not mean the rollback is complete. Retain policies, deletion protection, snapshots, log groups, network interfaces, elastic IPs, secret replicas, KMS deletion schedules and cross-account copies all have to be verified separately.
Closure criteria: every test recorded with a verdict; coverage measured for all accounts, regions and global services; reproducible evidence for every FAIL obtained without using real data; the reason and visibility impact written down for every BLOCKED; the absence of canary and temporary resources verified through inventory and audit; no unexpected cost or service impact; evidence masked and given a retention period; an owner and a target date assigned to each finding; short-term compensating measures defined for critical paths; retest criteria stated as an API decision or a detection outcome.
Closing
A good AWS penetration test earns its value not by laying hundreds of configuration results side by side, but by showing how those results combine from an attacker’s point of view. The most useful output states clearly which policy, trust, service role, event and data path can take a principal from initial access to higher impact, at which link the chain can be safely stopped, and how the fix will be measured. Because in the end what we test is not the definition of a permission; it is its decision.
This playbook is the framework of the approach the Red in Pulse team follows in AWS security assessments: a methodology that measures effective decisions rather than declared permissions, safely verifies attack paths with canary and dry-run discipline, and leaves the environment and its cost as clean as it started. If you would like the real effective permission boundaries of your environment - not merely the existence of its policies - tested with this rigor, you can contact us for a scoping conversation.
Concepts and abbreviations in this article
The AWS service for managing permissions of human and workload identities that access resources. Policies, roles, and conditions jointly determine authorization decisions.
A policy type that limits the maximum permissions available in AWS Organizations member accounts. An SCP does not grant permissions by itself.
The permissions an identity can actually use after all allow and deny layers have been evaluated.
A harmless, traceable test resource created to measure an access path without using real data.
The session-oriented version of the EC2 Instance Metadata Service. It reduces the risk of application-driven requests reaching instance-role credentials.