Skip to content

RINP // CYBERSECURITY SERVICES

Resource Center
Controlled Testing and RoE

Under which authorization, which boundary, andwhich escalation discipline is it carried out?

Offensive security engagements are carried out under written authorization, a defined scope, stop conditions, and a documented escalation path. The RoE framework clarifies the boundary of every testing step and the approval responsibility before the engagement begins.

// SERVICE-SPECIFIC CONTROLLED EXECUTION01

Service-specific controlled-execution frameworks

When you open each card, you see the authorization, scope, rules of engagement, approval gates, and escalation discipline of the corresponding service. The service itself opens on a separate page.

A seven-heading rules-of-engagement backbone that brings scope, authorization, boundaries, and accountability together in a single framework.

Controlled execution is built on written authorization, a defined scope, agreed rules of engagement, an agreed test window, approval gates, an established escalation line, and side-effect control. This card opens the seven-heading rules-of-engagement backbone of the Application Security Penetration Test service.

1. Written authorization and ownership verification

We take no technical step until we have verified in writing that the party acting as test owner holds authority over the asset. We require an authorization signature for every component in scope.

  • The authorization document carries an authorized signature, a scope definition, a validity period, and a revocation clause together.
  • Separate written permission is obtained for third-party components (CDN, external API, payment provider, and the like).
  • We provide advance notice consistent with cloud providers' permission policies (AWS, Azure, GCP).
  • A single-call withdrawal path exists to revoke authorization; it is suspended from one central point.

2. Scope statement: the Web, API, and Mobile boundary

The scope statement shows not only what is included but also what is excluded. Web, API, and Mobile components are defined as separate line items; the overlap discount is calculated once the scope statement is settled.

  • Web component: dynamic page range, number of roles, number of critical flows, and multi-tenant structure.
  • API component: endpoint range, REST/GraphQL/gRPC distinction, and authentication pattern.
  • Mobile component: number of platforms, binaries (APK/IPA), and the agreed mobile-to-backend flows.
  • Excluded areas: source code review, denial of service, third-party attacks, physical testing, and social engineering.

3. Rules of engagement (RoE): method, boundaries, and prohibitions

The rules of engagement fix in writing how the test will proceed and which behaviors are out of scope. The PTES Pre-engagement backbone and alignment with OWASP WSTG and the API Security Top 10 2023 form the technical foundation of these rules.

  • Denial of service (DoS), destructive stress, load testing, and real financial transactions are prohibited.
  • Social engineering and physical testing are out of scope in this package; there is a separate service line for that.
  • Automated scanning alone is not considered sufficient; manual validation and a safe proof of concept (PoC) are essential.
  • All request/response traces and PoC records are logged; unauthorized sharing is prohibited.
  • Data exfiltration attempts are performed only at the proof-of-intent level; actual data exfiltration is not carried out.

4. Test window and access model

The test window is fixed together with the choice of business hours or off-hours, the environment preference (staging or production-like), and the access model (white-box, gray-box, black-box). If the production environment is to be used, additional protection rules are put in writing.

  • Access model decision: white-box (full documentation and code access), gray-box (partial), or black-box.
  • When the production environment is chosen, a backup plan, a rollback procedure, and additional monitoring are required.
  • For an off-hours window, a two-sided on-call roster and a communication channel are settled in advance.
  • An IP allowlist, role-based test accounts, and API keys are provided together with the scope.

5. Approval gates: individual approval for sensitive steps

We do not pass certain steps automatically. We obtain written approval before database-level writes, state changes in a critical business flow, actions on payment endpoints, or attempts to read data on a per-user basis.

  • On write and update endpoints, an approval record is kept before the PoC.
  • No attempt is made to access the critical record set until the privilege escalation chain is complete.
  • In a multi-tenant environment, an explicitly owned approval is mandatory for attempts that cross the tenant boundary.

6. Escalation line: critical findings are reported immediately

We do not leave a critical-severity finding to the end of the report. The escalation line, contact person, backup contact, time-to-reach, and written notification channel are fixed in advance; a safe proof of concept is shared along with the notification.

  • Primary contact: technical decision-maker (CTO/Engineering Lead or Product Security Lead).
  • Backup contact: a single-call alternative person when the primary cannot be reached.
  • For a critical finding, written notification is delivered promptly over the agreed channel.
  • If signs of active exploitation are observed, the test is suspended and the situation is assessed jointly.

7. Side-effect control: no irreversible changes

The test does not produce an irreversible change. Deleting data, corrupting state, writing to real customer data, or executing real financial transactions is prohibited; all test data produced is marked in a traceable way and cleaned up at the end of the engagement.

  • Test-data label: all test records are created with a distinct prefix.
  • On endpoints that write to production data, we stop at the PoC level; full exploitation is not carried out.
  • Logging and alerting systems are informed in advance of the test window; false alarms are managed.
  • End-of-test cleanup: created users, records, and sessions are listed and closed out.

Written authorization, ownership verification, a defined scope, explicit approval gates, and a 24/7 escalation line; a promise of long-term stealth is not within the scope of this service.

This guide opens the rules-of-engagement backbone of the Network Security Penetration Test service across seven headings. What we assure is scope, methodological discipline, controlled execution, and delivery structure; we do not guarantee outcomes. The client's RoE addendum fills these seven headings with concrete details.

1. Written authorization and ownership: an asset with no established owner is not tested

We start every test engagement with written authorization and asset-ownership verification. Third-party assets do not enter scope until the authorization chain is complete. If an out-of-scope asset is discovered, we stop the test and notify the escalation channel.

  • The client-signed statement of work (SoW) and RoE addendum are received.
  • The asset-ownership list (IP, domain, SSID, segment) is verified.
  • If there is third-party infrastructure, additional authorization is obtained; otherwise it remains out of scope.

2. Scope statement: modules stay separate

The External Network, Internal Network, and Wireless Network modules are defined in separate columns in the scope statement. Each module's scope range, asset type, out-of-scope components, and module-specific constraints are fixed in the RoE addendum; scope expansion is managed through a written change request.

  • External Network: list of IPs, domains, subdomains, and VPN gateways.
  • Internal Network: host range, number of segments, and AD/Entra directory structure.
  • Wi-Fi: location, SSID, authentication type, and guest-versus-corporate separation.

3. RoE prohibitions: the natural boundaries of the service

The Network Security Penetration Test is carried out within a defined scope and a defined window. Some actions are out of the service's scope and are not performed; these are explicitly treated as prohibited in the rules of engagement.

  • DoS, stress testing, and load testing are not performed; service availability is not disrupted.
  • No persistence is left behind; no backdoor, scheduled task, or irreversible configuration is added.
  • No data deletion, modification of production data, or irreversible change is performed.
  • Long-term stealth, evading SOC/EDR defenses, and advancing undetected are not within this service's scope; when needed, the Modular Red Team Simulation framework is opened.

4. Test window and access model: varies by module

We fix the test window together with the choice of business hours or off-hours, the production freeze calendar, and the client's operational cadence. The access model varies by module; the Wi-Fi module is carried out on-site in most scenarios, and remote work covers only preparation and document review.

  • External Network: remote; black-box or limited gray-box.
  • Internal Network: remote, VPN, jumpbox, or on-site; with a gray-box account model.
  • Wi-Fi: on-site controlled assessment; no validation is performed without field access.

5. Approval gates: written approval for high-impact steps

Controlled privilege escalation, domain-administrator-level verification, an exploitation step against a production system, or PoC steps that touch a critical service pass through a written approval gate. The approval-gate request includes the expected impact, the rollback plan, and the escalation channel.

  • Privilege escalation approval: which asset, which method, and which rollback plan?
  • Lateral movement approval: target segment, expected footprint, and stop criterion.

6. Escalation line: a 24/7 emergency stop channel

A 24/7-reachable escalation line is fixed between the client and the test team. In the event of an unexpected impact, a production service outage, a third-party alarm, or a legal call, we stop the test immediately and deliver a structured notification to the client contact.

  • Primary and backup client contact (phone and email).
  • The test team's on-call line; the stop-time target is fixed in the RoE addendum.

7. Side-effect control: the production environment is protected

We plan steps against the production environment on a minimum-impact principle. The EDR/AV exception process is managed through a written instruction; exceptions are removed at the end of the test. Scan intensity, the number of concurrent connections, and the impact of exploitation tests are anticipated in advance from a side-effect-risk perspective.

  • EDR/AV exception process: on which machine, in which window, and who approved it?
  • Scan intensity: concurrency and rate limits are fixed in the RoE addendum.
  • Post-test cleanup: created accounts, sessions, and artifacts are removed.

A reference for reading ahead the written authorization, provider-policy compliance, scope statement, rules of engagement, test window, approval gates, escalation, and side-effect-control line.

This card is not a service page; it is the read-ahead surface of the controlled-execution discipline. It makes visible, across seven headings, what we commit to and which boundary we draw under the execution of the Cloud Security Penetration Test.

1. Written authorization and provider-policy compliance

We do not begin execution without written authorization. Provider policy (AWS Customer Support Policy for Penetration Testing, Microsoft Cloud Unified Pentest RoE, GCP Acceptable Use Policy) is an integral part of the authorization document.

  • AWS: no prior approval is required for permitted services; however, command-and-control simulation, testing against provider infrastructure, and actions that cause service disruption require written permission.
  • Microsoft Azure: prior approval has not been required since 2017; the client complies with the Microsoft Cloud Unified Penetration Testing Rules of Engagement and coordinated-notification principles.
  • Google Cloud: notifying Google is not mandatory; the client complies with the Cloud Platform Acceptable Use Policy and Terms of Service; the test must not affect other customers.
  • The authorization document fixes the scope, the account list, the test window, the source IP values, and the emergency contact; the contract carries a suspension gate.

2. Scope statement

The scope statement includes only the resources the client owns or is authorized over; provider infrastructure and the third-party tenant boundary remain outside.

  • The provider (AWS, Azure, GCP, or multi-cloud) along with the list of accounts, subscriptions, or projects is fixed in writing; count, boundary, and ownership are recorded exactly.
  • The region list, control-domain selection (IAM, network, storage, compute, Kubernetes, serverless, logging, data platform), and exceptions are written out explicitly.
  • CI/CD and IaC review are outside the core scope; unless taken as a separate add-on, they do not enter the statement. This distinction remains visible on the page as well.

3. RoE: prohibitions and boundaries

The rules of engagement explicitly prohibit destructive actions, scaling into the control plane, and impact on third-party SaaS; persistence outside the agreement and data disclosure are prohibited.

  • Service disruption (DoS, DDoS, flooding, stress testing) is not performed; finding validation is limited to the initial confirmation, and no action beyond exploitation is taken.
  • Scaling into the control plane is prohibited: no exploitation attempt is made that would carry a single account's control plane into the provider's shared infrastructure.
  • No action is taken that would affect third-party SaaS, a tenant boundary, a subscription, or a project; out-of-scope resources are tracked under recording at every stage.
  • No destructive structural change is made; deleting data, removing permissions, deleting records, leaving a persistent backdoor, and billing abuse are prohibited.

4. Test window and access model

We write the test window taking business load and operational load into account; the access model limits authority to what is necessary.

  • The default approach is gray-box. A read-only IAM role is sufficient for most scopes; when attack-path quality requires it, a controlled test role is used.
  • The test window is fixed together with the start and end times, region, MFA, VPN, jump server, PIM, and JIT constraints.
  • A production-only environment, a narrow window, and rollback uncertainty are price and scope drivers; the rollback line is defined in advance.

5. Approval gates

Controlled exploitation in a production account proceeds with an approval gate at each step; actions near the control plane and data-read attempts proceed under explicit approval.

  • A high-impact action (reading private data, assuming a production role, attempting cross-tenant movement) is not performed without explicit approval from the client's authorized person.
  • The approval chain is recorded in writing; the moment of approval, who approved, and what was approved are recorded; the contract's suspension gate remains open at every stage.

6. Escalation

If a side-effect risk is observed during attack-path work or provider-quota API limits are triggered, we stop the work immediately and call the escalation line.

  • The emergency-stop contact (on the client side and the Red in Pulse side) is fixed in the authorization document; an off-hours access line is defined.
  • Provider-quota API call limits (for example AWS service quotas, Azure throttling, GCP quota) are observed; on an approaching signal, execution speed is reduced or halted.
  • If a report must be conveyed to the provider, coordinated-notification principles are applied (MSRC for Microsoft, the Vulnerability Reward Program for Google, the Security abuse line for AWS).

7. Side-effect control

The CloudTrail, Activity Log, and Cloud Audit Logs recording requirement remains on throughout the test; billing and quota alerts are monitored.

  • The AWS CloudTrail, Azure Activity Log, and Google Cloud Audit Logs recording stream is kept on without interruption; if a loss of monitoring is detected, the work is stopped.
  • Billing alerts are kept on; if an exploitation attempt leads to unusual resource consumption, the client side is notified immediately.
  • The post-test cleanup step is written: created test users, roles, keys, and temporary resources are decommissioned under recording.

Written authorization, channel policy, a just-culture framework, and KVKK (Turkish data protection law) compliance are the four pillars of running a Social Engineering Simulation.

The Social Engineering Simulation runs under seven headings. Each heading makes the ethical boundary, the legal framework, and the control over employee impact visible.

1. Written authorization and triple approval

Before the campaign, we obtain the written approval of the HR, legal, and Information Security units. The processing of employee personal data under KVKK, a disclosure approach balanced against campaign integrity, and data minimization are settled at this step. The campaign is not launched until the approval chain is complete.

  • HR approval: employee impact and communication policy.
  • Legal approval: KVKK compliance and disclosure approach.
  • Information Security approval: technical scope and escalation line.

2. Scope statement: channel, wave, and scenario family

The scope statement clarifies the channel selection, the number of waves, and the scenario family. Email and QR are the core channels; SMS phishing and voice phishing are activated as a pilot or add-on under a separate approval. Each channel's target audience, language, and content-approval flow remain within the written scope.

  • Core channel: email and QR.
  • Pilot channel: SMS phishing, a separate approval chain.
  • Pilot channel: voice phishing, a separate approval chain.

3. Execution rules and just culture

The campaign is harmless and controlled. Real financial transactions, real credential abuse, persistence, physical intrusion, and third-party interaction are out of scope. Harassing language, shaming framing, and penalizing employees are prohibited. Our framework focuses not on the question "who was fooled?" but on "at which point is the process breaking down?"

  • Real financial transactions are prohibited.
  • Real credential abuse is prohibited.
  • Harassing language and penalizing individuals are prohibited.
  • The framework is process-focused; it is not blame-oriented.

4. Test window and period constraints

We set the test window in writing as business hours or off-hours; busy periods are closed to the campaign. Religious holidays, public holidays, financial year-end, crisis-communication periods, layoff processes, and merger/acquisition periods are automatically out of scope; during these periods we conduct a separate risk assessment.

  • Overlapping with religious and public holidays is prohibited.
  • Layoff and merger periods are automatically closed.
  • Crisis-communication period: a separate risk assessment.

5. Approval gates and executive campaigns

Campaigns aimed at senior management, finance leaders, and privileged users pass through an additional approval gate. These audiences are not targeted in a mixed way within a single wave; the scope, content, and distribution order are settled through a separate approval chain. Approval gates are preserved in writing at every critical step of the campaign.

  • Senior-management campaigns: an additional approval gate.
  • Targeting finance, senior management, and privileged users mixed within a single wave is prohibited.

6. Escalation and emergency stop

If a real security incident is triggered during the campaign, an employee shows signs of psychological harm, or a risk of operational disruption arises, we stop the campaign immediately. The escalation line is defined in writing; the contact person, stop authority, and restart conditions are fixed in the RoE file.

  • If a real incident is triggered: immediate stop.
  • A signal of psychological harm in an employee: immediate stop.
  • Risk of operational disruption: immediate stop.

7. Side-effect control and KVKK compliance

Employee psychology is an integral input to campaign design. In line with the data-minimization principle of KVKK Art. 4, we process only the personal data needed for measurement; after the campaign we apply anonymized reporting. Individual names and identities are kept out of the report; the distribution by segment, persona, and wave is preserved.

  • KVKK Art. 4: the data-minimization principle.
  • Anonymized reporting; no individual names appear in the report.
  • Post-campaign training access is open to all targets.

A governance framework built on a minimum-three-month program SoW, a separate RoE window per cadence, and program-fatigue protection.

Continuous Penetration Testing is not a single-window test; it is a program that runs on a monthly cadence, is re-approved at each cadence, and is executed without straining the release tempo. This card opens the seven headings of that governance framework.

1. Written authorization and the program SoW

The program begins with the client's written authorization and a minimum-three-month Program SoW (Statement of Work). The SoW records the duration, the target asset set, the number of cadences, the validation rights, and the out-of-scope items. No cadence is opened without an authorization document.

  • Written authorization signed by the client's authorized person and contract approval are mandatory.
  • The Program SoW duration is a minimum of 3 months; requests for a shorter period are directed to the Application Security Penetration Test page.
  • The SoW scope separates the target assets, the monthly cadence entitlement, the deep-review entitlement, and the out-of-cadence release-validation entitlement.

2. Scope statement

The scope statement records, in fixed language, which target assets will be tested throughout the cadence, which roles and critical flows are in scope, and which surfaces remain out of scope. Scope is decided at the contract stage, not at the start of the month.

  • Target asset types: web application and API surface (mobile, internal network, and cloud posture are out of standard scope).
  • The number of monthly cadence entitlements (Basic: 1, Deep Review: 2, Advanced: 3) and the deep-review entitlement are recorded according to the package tier.
  • Out of scope: 24/7 monitoring, incident response, BAS, Purple Team, Red Team, social engineering, DDoS, and destructive testing.

3. A RoE window per cadence

We open a separate RoE window for each monthly cadence; this window does not behave like a cumulative authorization. Actions authorized in the previous cadence are not automatically carried into the next. Surprise scope expansion and testing of assets outside the release are prohibited.

  • No cumulative authorization: each cadence is opened, closed, and recorded with its own RoE window.
  • No surprise scope expansion: an asset or role outside the SoW is not unilaterally added to that cadence.
  • No out-of-release assets: testing is carried out only on the target assets defined in the SoW.
  • A risky action (privilege escalation, data write, session tampering) requires an additional written approval.

4. Test window and access model

The test window is placed on the monthly calendar; release-freeze days and critical business days are not included in the window. Test accounts, OpenAPI or Postman collections, and release-notification integration are ready at the start of the month. If release priority conflicts, we shift the window.

  • The monthly calendar is approved between the client and the test team at the start of the month.
  • Release-freeze or critical business days are left outside the window; if priorities conflict, the window is shifted.
  • The test account, OpenAPI/Postman collection, and release-notification flow are verified in each cadence.

5. Approval gates

Detecting a critical finding, opening the remediation window, and re-validation are each a separate approval gate. The client's authorized person gives written approval at each gate; no step advances until the gate is passed.

  • Gate 1: Critical-finding notification: written notification to the client's authorized person within 24 hours and a validation call.
  • Gate 2: Remediation window: the client team presents the remediation plan; a re-validation date is set.
  • Gate 3: Re-validation approval: authorization for a controlled retest is written separately; evidence that the finding is remediated is recorded.

6. Escalation and reporting

At the end of each cadence, we deliver an executive summary and a technical findings list. When a critical finding is detected, the notification flow operates within 24 hours; if a risk of production impact arises, the work is stopped and the escalation line is activated.

  • End-of-cadence report: executive summary, technical findings list, remediation rate, and the priority area for the next month.
  • A 24-hour notification for a critical finding; a 72-hour notification line for a high finding.
  • If a risk of production impact arises, the work is stopped; the technical contact and emergency contact are called.

7. Side-effect control and program-fatigue protection

We plan the monthly cadence so it does not strain the production tempo; alert cacophony (a pile of unnecessary warnings), a downpour of low-priority findings, and overwhelming the product team are prevented. We report only findings that are prioritized by business impact and validated.

  • No alert cacophony: instead of individual informational-level notifications, a consolidated report at the end of the month.
  • Production-tempo alignment: freeze days are outside testing; no risky action is taken during critical business hours.
  • Instead of a pile of low-priority findings, a findings list filtered by business impact and validated.

Engagement authorization, scope statement, rules of engagement, test window, approval gates, escalation line, and side-effect control are fixed as a single document backbone.

The Modular Red Team Simulation holds covert scenarios that run for weeks or months. Controlled execution carries together senior-management approval, narrow stakeholder briefing, ethical boundary windows, and an escalation line that prevents confusion with real incident response.

Framework: TIBER-EU · CBEST · AASE

1. Written authorization and senior-management approval: the authorization chain

The engagement begins with an authorization document signed by senior management (on the CISO or CTO axis, and with the CEO informed in critical scenarios). The document fixes the scope, the starting assumption, the ethical boundaries, the record-keeping principle, and the confidentiality plane of the contract. A narrow set of stakeholders (white team) is designated; the blue team and the broader team are not informed.

  • Senior-management-signed authorization document: the CISO or CTO axis; the CEO informed in critical scenarios.
  • White team list: 2-4 people; CISO, security director, senior-management liaison, and legal.
  • The blue team is out of scope; so that genuine detection capability can be measured, execution is kept covert from the blue team.

2. Scope statement: module selection and the scope line

The scope statement fixes the core module and the extensions separately: RT-1 threat-actor behavior, RT-2 objective-driven critical-asset chain, RT-3 assumed-breach lateral movement, RT-4A human-layer initial access. The crown jewel and the starting assumption are defined in writing; out-of-scope assets (third parties, regulators, partners) are listed separately.

  • Core module: RT-1 Threat Actor Simulation, RT-2 Objective-Driven Red Team, or RT-3 Assumed Breach.
  • Extension: RT-4A human-layer initial access, RT-4B physical entry, RT-5 purple/BAS validation.
  • The critical asset and starting assumption are written; out-of-scope assets are listed separately.

3. Rules of engagement (RoE): ethical boundary windows

The rules of engagement fix which actions the work does not fall under. Destructive actions, planting persistent malware, deploying ransomware, exfiltrating real production data, and unauthorized third-party interaction are out of scope by default. In scenarios that require data access, a representative-data label is used; no real customer record appears in the evidence package.

  • No destructive actions: DoS, capacity testing, or execution that impairs system integrity is not performed.
  • No deploying ransomware or real malware; only controlled emulation tools are used.
  • No real data exfiltration: the evidence package carries a representative-data label; no real customer data leaves.
  • No third-party supplier assets: an unauthorized partner, regulator, or service-provider system is out of scope.
  • Ethical boundary windows: critical production hours, seasonal peak periods, and audit calendars are written down.

4. Test window: duration, cadence, and white team coordination

Unlike the defined window of the Network Security Penetration Test service, the Modular Red Team Simulation carries an execution duration ranging from weeks to months. Depending on scenario complexity, 2-6 weeks is the typical range; for engagements tied to a regulated framework (TIBER-EU/CBEST), the duration is longer. The white team coordinates through a daily or weekly coordination channel; every critical step is recorded.

  • The typical duration is 2-6 weeks; longer under a regulated framework.
  • A daily or weekly coordination channel for the white team; an activity summary is reported.
  • Access model: assisted, gray-box, or black-box; the written assumption is fixed.

5. Approval gates: trusted-agent approval at critical steps

Critical steps on the attack path (impact on a production system, privilege escalation, persistence emulation, controlled objective impact) pass through trusted-agent approval before execution. The approval is recorded and can be withdrawn. This discipline both protects senior management from side-effect surprises and keeps the evidence package auditable.

6. Escalation line: preventing confusion with real incident response

If a SOC or MSSP mistakes the simulation activity for a real attack and triggers the incident-response process, the trusted-agent channel is activated and gives the deconfliction signal confirming that the activity is controlled. The same line also ensures that, when a real attack arrives concurrently, the simulation is suspended immediately and resources are released to the real incident response.

7. Side-effect control: deconfliction, recording, and cleanup

Side-effect control runs along three lines: deconfliction (a time and technique record so the activity is not confused with a real threat actor), activity logging (recording each technical step with a timestamp, target, and outcome), and footprint cleanup (accounts, files, agents, and persistence artifacts created during execution are cleaned up at the end of the contract; the cleanup is recorded as well).

  • Deconfliction: activity time and technique record; the risk of confusion with a real actor is managed.
  • Activity logging: each step is recorded with a timestamp, target, and outcome.
  • Footprint cleanup: created accounts, files, agents, and persistence artifacts are cleaned up at the end of the contract.

Controlled execution under written authorization, a scope statement, third-party model-provider terms, approval gates, and side-effect control.

The Generative AI Red Team is not a random jailbreak exercise; it is carried out with written rules over the agent workflow, RAG sources, MCP servers, third-party model-provider terms of use, and write-capable tools. This card opens the seven headings of those rules.

1. Written authorization

The engagement begins only with written authorization, verification of agent and account ownership, and a preliminary check of the terms of use of the integrated third-party model providers. The provider infrastructure itself is out of scope; the scope is the client's agent workflow.

  • A signed statement of work (SoW) and the approval of the legal owner on the client side.
  • Written verification of ownership of the agent workflow, accounts, RAG, and MCP.
  • A check of the third-party model providers' acceptable-use policies (AUP) and predetermining a compliant test profile.
  • Provider infrastructure is out of scope; no black-box attack testing is performed against it.

2. Scope statement

The scope is defined at the unit level of the agent workflow. The tool library, RAG sources, MCP servers, system prompts, and the role/approval boundary are listed in scope; components that remain out of scope are written out separately.

  • In scope: the agent workflow, tools/connectors, RAG indexes, MCP servers, the system prompt, and the safe-execution boundary.
  • In scope: the authentication, authorization, approval, and token-management layers.
  • Out of scope: model-provider infrastructure, model-file provenance, social engineering, physical testing, and third-party systems.
  • Out of scope: DoS, load, and stress testing; uncontrolled data exfiltration on production.

3. Rules of engagement (RoE)

The rules of engagement fix in writing the permitted and prohibited actions, the sensitivity boundaries, and the stop criteria. The test goes deep enough to prove the runtime chain; it does not touch the production user's data or real actions.

  • Permitted: indirect prompt injection, jailbreak, sensitive-information disclosure testing, tool abuse, and MCP authorization-abuse attempts.
  • Prohibited: reaching production user data; synthetic or masked data is used instead.
  • Prohibited: actions that trigger real financial transactions, payments, or email/message sending; a sandbox equivalent is used instead.
  • Prohibited: permanently altering model behavior, fine-tuning, or performing memory manipulation in the production environment.
  • Stop criterion: suspicion of a customer-data leak, impact on a production user, or a signal of a provider AUP violation.

4. Test window and access model

We run the test primarily on a sandbox or canary tenant. The production scope is handled only through read-only or strictly limited scenarios, under written agreement, and within a narrow window.

  • Sandbox: a production-equivalent structure of the agent; synthetic user data, mock connectors, and isolated MCP servers.
  • Canary tenant: production architecture and an isolated account boundary; a controlled PoC space separate from real customer data.
  • Production read-only: only read-capable tools; write tools are behind an approval gate.
  • The test window is not made to overlap with release periods and the model-provider update calendar.

5. Approval gates

High-impact scenarios proceed behind approval gates. Scope and RoE approval at the start; high-impact actions, production touches, and report acceptance are recorded at separate gates.

  • Scope and RoE approval (kickoff): the signatures of the agent-workflow owner, the security sponsor, and the legal owner.
  • Written approval before high-risk prompt families and write-capable tool calls.
  • Approval before a production touch: which tool, which scope, and which window.
  • In the reporting phase, a finding-acceptance session and retest-scope approval.

6. Escalation

Runtime testing carries the risk of triggering model-provider terms of use. The escalation line is defined on both sides for instant coordination in the event of a user-data suspicion or production impact.

  • Client-side escalation: the agent owner, security sponsor, and legal counsel contact (24 hours).
  • Service-provider-side escalation: the test team lead, project manager, and security officer.
  • A model-provider AUP-violation signal: the test is stopped, a record is taken, and coordination with the provider is carried out.

7. Side-effect control

When the engagement ends, the test data, the agent log history, and the provider-side conversation records are handled through a written cleanup procedure. We frame customer data leaving no trace from the test as an element of trust.

  • The agent log history created during the test is anonymized or deleted.
  • A retention policy for the provider-side conversation records is determined in advance.
  • The test data (synthetic or masked) is destroyed at the end of the engagement.
  • The findings report's retention period and access permission are written; if a bilingual option exists, it is stated.

Controlled-execution framework: seven headings; authorization, scope, rules, schedule, approval, escalation, and side-effect control.

AI Model Supply Chain Assurance is not an attack simulation. We assess the model pipeline, the registry, and the release chain under authorized access, a read-only rule, and boundaries that do not harm production. We write the seven headings together; a missing heading does not let the engagement start.

1. Written authorization

We do not begin the assessment without the model owner's written authorization. In the authorization document, the model inventory, the registry access level, the training-data inventory read limits, and the responsible stakeholders are listed by name.

If a third-party base model is used, the provider terms of use (the Hugging Face model license, closed-provider API terms of use, fine-tuning license terms) are assessed as an addendum to the authorization document. License interpretation is not a legal opinion; it stays at the flagging level.

2. Scope statement

The assessment scope is defined at the level of the model version list, the training/fine-tune pipeline, the registry, and the release targets (staging, production, multi-tenant, multi-region). Areas that remain out of scope are written out explicitly.

Out of scope: the runtime prompt, tool, RAG, and MCP attack chain belongs to the Generative AI Red Team assessment; the internal offensive workflow falls under AI for Penetration Testing and Red Team. Legal license opinions and accuracy and fairness validation are outside the service scope.

3. Controlled-execution rules

Access to the registry, the pipeline, and the release manifests is read-only by default. We do not perform unauthorized promote, rollback, or configuration changes; write access is applied only with a separate written approval and within a recorded window.

Access to the training dataset is taken through sampling and under the data-minimization principle. Sensitive data (special-category personal data, trade secrets, customer data) is masked or hashed; the original data is not exfiltrated.

We do not perform destructive testing on real model artifacts; the live model's weights are not altered. Disrupting the production release pipeline is prohibited; if a finding is to be validated during production, it is carried out under written approval and a controlled window.

4. Test window and access model

We build the access model incrementally according to the evidence mode. Reading the CI/CD pipeline, reading the model registry, sample access to the training-data inventory, and reading the attestation log are the primary links. In air-gapped or regulated environments, an on-site engagement window is opened.

We plan the engagement window so it does not overlap with the MLOps calendar's retraining, fine-tune, and release waves. The emergency exit stays in writing at each side's contact point; the communication line is bound to a 24-hour response threshold.

5. Approval gates

If the training data contains a sensitive category, the non-disclosure agreement (NDA) is reinforced and the data owner (governance or legal) is bound to an additional approval gate. If there is special-category personal data, a separate approval is obtained under KVKK Art. 6.

If a fine-tune has been done on a third-party base model, the provider terms of use are added to the approval chain. Multiple stakeholders (ML platform, security, governance, and legal) sign at a joint approval gate; a missing signature stops the engagement.

6. Escalation line

If a missing signature in the registry, a break in the chain of evidence, or a trace of an unauthorized promote is detected, we immediately inform the primary technical contact and the security sponsor. A preliminary report is shared within 24 hours; the remediation plan is tied to a hardening backlog.

We escalate a finding with possible sensitive-data leakage, a sign of a third-party license violation, or regulatory impact directly to management-level stakeholders. The information-sharing path, the notification line, and the written record are clarified from the outset.

7. Side-effect control

Personal data in the training data, IP-protected content, and regulated data are defined from the outset. A separate protection rule is applied for KVKK Art. 6 sensitive data and special-category personal data; data minimization is the default principle.

Customer data is not exfiltrated; samples are masked and deleted at the end of the contract term. The default retention period is 30 days, or it is set differently by contract. Audit logs are kept in a separate channel for the duration of the assessment.

Written authorization, use-case scope, sandbox, canary, and limited-pilot discipline with human approval at every stage; a fully autonomous attack agent is not within the scope of this service.

This guide opens the rules-of-engagement backbone of the AI for Penetration Testing and Red Team internal pilot across seven headings. What we assure is scope, methodological discipline, controlled execution, and delivery structure; we do not guarantee outcomes. The internal pilot's RoE addendum fills these seven headings with concrete details.

Seven-heading RoE backboneInternal leverage line

1. Written authorization and pilot ownership: a scenario with an undefined owner is not piloted

We start every internal pilot engagement with written authorization and the commitment of the internal offensive lead. Use cases with no established owner or without written acceptance criteria are not taken into the pilot scope. The pilot owner is the internal offensive lead; the final finding decision rests with the expert analyst.

  • The pilot engagement authorization signed by the internal offensive lead and the RoE addendum are received.
  • The use-case list, its owner, and the acceptance criteria are fixed in writing.
  • The third-party AI vendor's AUP and data-processing terms are checked.

2. Scope statement: a separate owner and a separate eval set for each scenario

The pilot scope is defined as a written list of offensive-workflow use cases: reconnaissance, hypothesis, evidence processing, reporting, and retest. Each scenario carries a separate owner, separate acceptance criteria, and a separate eval set; "abstract AI use" is not a scope. Scope expansion is managed through a written change request.

  • Reconnaissance / target pre-assessment workflow: with its owner and acceptance criteria.
  • Test-hypothesis workflow: ownership of the finding-pattern library.
  • Evidence-to-finding and retest workflow: ownership of the report template.
  • Multi-step orchestration: a separate RoE addendum is required for an advanced pilot.

3. RoE prohibitions: the natural boundaries of the service

The internal pilot is carried out within a defined scope and a human-supervised workflow. Some actions are out of the service's scope and are not performed; these are explicitly treated as prohibited in the rules of engagement. Autonomous production actions and any promise of work that bypasses the expert are refused.

  • No fully autonomous attack agent is built; there is human approval for every action.
  • In production security testing, automation is enabled only with written approval at every stage.
  • A workflow outside the eval-set scope is not moved to production; each scenario passes an acceptance gate.
  • No use contrary to the third-party model's AUP is made; the license and data-processing terms are preserved.
  • Attack testing of the client's AI product is outside this scope; when needed, the Generative AI Red Team is opened.

4. Test window and access model: sandbox, canary, and limited pilot

The internal pilot begins in a sandbox or an isolated development environment; it moves through a canary run to a limited real scenario; the final pilot scope expands through a written acceptance gate. The move to the production environment is not automatic; every scale-up is bound to a new approval gate.

  • Sandbox: an isolated development environment; there is no contact with real customer data.
  • Canary: anonymized sample artifacts and a narrow scenario.
  • Limited pilot: a use-case set expanded through written acceptance.

5. Approval gates: human approval at every stage

The use-case selection, RoE, artifact and access, action level, PoC-acceptance, and pilot-acceptance gates are opened with written human approval. Automatic-action boundaries are fixed in a least-privilege action matrix; an action outside the matrix is not carried out without approval. The final finding decision rests with the expert analyst.

  • Use-case approval: scope, owner, and acceptance criteria in writing.
  • Action-level approval: the distinction between read-only, draft generation, and execution is kept clear.
  • PoC and pilot-acceptance gates: eval-set results are an input to the gate decision.

6. Escalation line: a drop in the eval set is a trigger

A structured escalation line is fixed between the internal offensive lead and the pilot team. The eval set dropping below the acceptance threshold, an unexpected action, a third-party model warning, or a side-effect signal is a trigger for stopping the pilot immediately and for a structured notification.

  • Primary and backup internal offensive contact (phone and email).
  • The eval-set drop threshold is written; when it is exceeded, the pilot stops and is revised.

7. Side-effect control: customer data is not fed to the model

Within the pilot, we do not feed real customer data into the AI model as a training input. Anonymization and data minimization are applied from the outset; the log-cleanup process is written. The AI output does not replace the expert's decision; every critical decision is validated by the expert analyst.

The definition and decision impact of each service’s primary evidence object.

The prerequisites, access and stakeholders you need to prepare before testing begins.

It opens in which situation which service is the right start.

Let’s determine the target surface, the authority arrangement and the test window together in a discovery call; binding scope is fixed with an approved scope-of-work document.

  • Training-data policy: customer evidence is not fed to the model; opt-out is verified.
  • Anonymization: personal data and customer identity are removed from the outset.
  • Log cleanup: post-pilot artifacts and sessions are removed through a written process.
  • Replacing the expert's decision is prohibited; the AI output is a draft, not a decision.

// CONTROLLED EXECUTION

Let's clarify the scope and rules of engagement together.

Let's determine the target surface, the authorization arrangement, and the test window together in a discovery call; the binding scope is finalized with an approved scope-of-work document.