RINP // CYBERSECURITY SERVICES
Resource CenterIn each service,what does "verified" mean?
For each offensive security service, we explain which verified security finding it produces, how that finding is recorded, and which business decision it supports.
Service-based evidence object reference cards
Open each card to see the definition, production discipline, and decision impact of the primary evidence object that service produces. The service itself opens on a separate page.
A verified security finding produced within a critical business workflow, plus an initial remediation list ordered by business impact.
Application Security Penetration Testing produces an output that goes well beyond a one-line note: a reproducible security finding within a critical business workflow, a remediation list ordered by business impact and risk level, and a readable decision file. This card opens up the concrete components of that file.
Evidence object definition
The evidence object is a security finding demonstrated reproducibly across the agreed critical user journey, together with visibility into the affected privilege, tenant boundary, and business logic, plus a remediation list ordered by business impact. Beyond a single finding report, it carries a decision-ready evidence file.
- Reproducible evidence: a safe PoC, request/response trace, screen recording, or command sequence.
- Impact map: a clear explanation framed around privilege boundary, tenant separation, business logic finding, or data leakage.
- Prioritization: a meaningful ranking combining CVSS v3.1, EPSS, and business impact. This goes beyond a plain list.
- Dual-Layer delivery: a risk picture for leadership and an owner-assigned work list for the technical team.
Production steps (PTES and OWASP aligned)
The evidence object is produced in a flow aligned with the seven PTES phases and the OWASP WSTG, OWASP API Security Top 10 2023, and OWASP MASVS/MASTG frameworks. The output is a verification pipeline, not merely a document.
- Scope and rules of engagement verification (PTES Pre-engagement).
- Reconnaissance and threat-focused preparation (PTES Intelligence and Threat Modeling).
- Deep manual testing and exploitation verification (PTES Vulnerability Analysis and Exploitation).
- Delivery of the executive summary, technical report, PoC production, and prioritized remediation list (PTES Reporting).
- Optional remediation verification and retesting (the remediate-and-verify loop).
Reading example: anonymized finding structure
The structure below shows how a single finding record from an Application Security Penetration Testing report is read. It contains no client, sector, or product name. Only the reading order carries the discipline of the evidence object.
Cross-tenant privilege leakage: bypassing role checks in the multi-tenant model of a managed relational database.
Read access to another tenant's critical record set within the same release. A business logic finding.
Request/response trace, safe PoC, and reproduction steps (repeated across three separate sessions).
Critical severity. The first item on the prioritized remediation list, ranked first in the pre-release remediation order.
Decision impact: turning into a work list
The prioritized remediation list is more than a single table. Each item moves into the technical team's work tracking with a defined owner, priority, and verification logic. The output enters the decision set of the next release.
- For each item, an owner-assigned responsibility, a priority label, and a re-verification note.
- Field discipline that can be exported to Jira, GitHub Issues, or similar systems.
- A risk view mapped to the executive summary. The decision picture leadership wants to see.
Neighboring evidence objects
When the decision priority differs, the evidence object differs as well. The distinction below directs the reader to the right service.
- Continuous Penetration Testing: cadence-based verification, measurable remediation, and management visibility.
- Cloud Security Penetration Testing: privilege escalation chain, data access path, and transitive privilege visibility.
Beyond scan output: a verified security finding, a chainable access flaw, segmentation impact, and a prioritized remediation track.
The output of Network Security Penetration Testing goes beyond a finding list. What is produced is the distillation, through controlled verification, of service-based findings on the external network, identity and segmentation flaws on the internal network, and lateral movement risk at the wireless layer, all captured into an evidence file.
Evidence object definition: it varies by module
External Network: technical impact proven through manual verification on internet-exposed services: a misconfigured edge layer, an externally facing management interface, weak authentication, and a verified service finding. The output shows which service is exploitable and to what degree.
Internal Network: verified access impact via identity, segmentation, and critical services: privilege surface in the Active Directory or Entra ID flow, service account weaknesses, and bypassing segment separation through a chainable access flaw. A lateral movement narrative is added for Package Level 2 and above.
Wireless Network (Wi-Fi): verification output tied to the SSID, authentication mechanism, and segmentation separation: corporate versus guest SSID separation, WPA2/WPA3 and 802.1X flow, rogue access point risk, and the possibility of pivoting from guest to internal network. For Package Level 2 and above, the presence or evidenced absence of lateral movement is produced as a result.
How the evidence is produced
The production flow follows the four-phase discipline of NIST SP 800-115: planning and authorization, discovery, attack simulation, and reporting. The PTES Vulnerability Analysis and Exploitation sub-phases are carried out with controlled manual verification. Each finding is recorded with safe reproduction steps, an impact definition, and an evidence artifact.
- Planning: written authorization, asset ownership, RoE, test window, and escalation path are fixed.
- Discovery: passive and active inventory. Service, version, certificate, and configuration signals are collected.
- Attack simulation: controlled manual verification after vulnerability analysis. Side-effect monitoring is kept active.
- Reporting: the executive summary, technical report, evidence file, and prioritized remediation order are delivered together.
Reading example: anonymized case
Internal network module at a high-volume services-sector organization. By chaining the weak password of an old maintenance account, a service account ticket obtained via Kerberoasting, and a misconfigured trust relationship, we documented with an evidence artifact a privilege escalation path reaching domain administrator level. Because its segmentation impact was high, the finding landed in the top three of the prioritized remediation order.
This reading example is structural. The sector category, product name, individuals, and geographic identifiers have been anonymized. The evidence file template is expanded on the Evidence Package Examples page.
Decision impact: how the evidence reaches the decision table
When the evidence object becomes a prioritized remediation order, each item's risk level is set by computing business impact, exploit value, and segmentation impact together. The executive summary shows the action track. The technical report distills it into an owner-clear work list. Dual-Layer delivery carries the same evidence to two different decision tables with the same clarity.
- Business impact: which asset or business process is at risk?
- Exploit value: how exploitable is the flaw, and how far does it chain?
- Segmentation impact: when remediated, which adjacent risks close alongside it?
Neighboring evidence objects: evidence that looks close but stands apart
In Modular Red Team Simulation the evidence object is the attack path: the chain to a critical target, the detection gap matrix, and the control that breaks the chain. In Network Security Penetration Testing the evidence is a network surface finding and carries no promise of an attack path. The two evidence objects are related but not synonymous.
In Cloud Security Penetration Testing the evidence object is the privilege escalation chain and the data access path. AWS, Azure, and GCP control-plane verification is specific to the cloud layer. The internal network component can combine with Network Security Penetration Testing in hybrid identity flows. However, ownership of the cloud control plane is the subject of Cloud Security Penetration Testing.
The primary evidence object has three layers: the privilege escalation chain, the data access path, and transitive privilege visibility.
This card is more than a service page. It is a pre-decision reading surface. Here you see the structure of the evidence you will receive after Cloud Security Penetration Testing, its production steps, and which decision it opens up.
The primary evidence object: what do we produce?
The primary evidence object of Cloud Security Penetration Testing has three layers: the privilege escalation chain, the data access path, and transitive privilege visibility. Together, the three feed the prioritized remediation order based on business impact.
- Privilege escalation chain: a set of privilege relationships that look trivial in isolation but together escalate privilege (AWS IAM roles and trust policy, subscription privilege delegation via Azure Entra ID, project-level delegation via GCP service accounts).
- Data access path: the verified ability to read or modify data at the end of the privilege chain (S3, Blob, Cloud Storage; RDS, SQL, BigQuery; Secrets Manager, Key Vault, Secret Manager).
- Transitive privilege visibility: making visible, through runtime verification rather than a records panel, the privileges that are not granted directly but inherited through one or more roles, policies, or groups.
How the evidence is produced
Production begins under written authorization consistent with the provider's execution rules. Inventory is taken through CIS Benchmark-aligned discovery. The exploitation path is prioritized using the MITRE ATT&CK Cloud technique map.
- The account, subscription, or project scope and the provider rules (AWS Customer Support Policy, Microsoft Cloud Unified Pentest RoE, GCP Acceptable Use Policy) are clarified before written authorization.
- The configuration baseline is read through discovery aligned with the CIS Foundations Benchmarks (AWS, Azure, GCP) and provider security benchmarks (including the Microsoft Cloud Security Benchmark). A deviation list alone does not count as evidence.
- The exploitation path is mapped within the MITRE ATT&CK Cloud framework: T1098.003 additional cloud roles, T1078.004 valid cloud accounts, cross-tenant role assumption, and transitive delegation techniques are exercised through controlled verification.
- Verification begins with a gray-box default. Attack-path quality is increased with a controlled test role where needed. Read-only visibility is sufficient for most scopes. No destructive action is taken.
Reading example: anonymized evidence chain
The narrative below is structural. The identity of a real client engagement is concealed while the evidence structure is preserved. The sector is referred to by a generic category and carries no numeric fingerprint.
A high-volume digital product team, a multi-account landing zone architecture (AWS landing zone), a shared control plane, and production/development separation.
An application role in the development account could assume a role into the shared control account through an overly broad trust policy. From there, a transitive role chain extending into the production account was verified.
At the end of the chain, read access to the object store holding production data and to secrets management was confirmed with a controlled PoC. No destructive action was taken.
Two role delegations that could not be read from the IAM control panel were made visible through runtime verification. The root cause in the landing zone template was traced to a single point.
Which decision the evidence object opens up
The output of Cloud Security Penetration Testing is more than a finding list. It is the top 10 actions prioritized by business impact. The ranking logic combines business impact, exploit value, and segmentation impact to set the risk level of each action.
- The top 10 actions give the remediation order. They carry a list that is owner-clear for the technical team and that leadership can read defensibly.
- Re-verification logic is included for remediation visibility. A discounted retest can be scheduled within 45 days on the same scope and same architecture.
- The control-mapping summary is handed over to a defensible assurance layer for client security reviews or audit priorities (in a portable or audit-ready form).
Neighboring evidence objects
The cloud evidence object does not replace the network or application evidence object on its own. The three tracks remain distinct and complement one another.
- Network Security Penetration Testing: a verified security finding, a chainable access flaw, and segmentation impact. It is the evidence object of the network surface and does not cover the cloud layer.
- Application Security Penetration Testing: verified technical evidence in a critical flow and an initial remediation order. It is the evidence object of the application layer and does not cover the cloud layer.
- CSPM output: a configuration deviation list. It shows neither what was exploited, nor the segmentation impact, nor the transitive privilege chain. It does not count as an evidence object.
The evidence object is a five-metric human-layer measurement package. Click rate is context data, not the decision object.
Click rate is context data, not the decision object itself. The measurement package is built on the detect-and-report rate, time to report, correct escalation percentage, process breakdown, and control effectiveness.
Evidence object: the five-metric measurement package
The primary evidence object of the Social Engineering Simulation is not reduced to a single rate. Five metrics are read together, and the clear decision emerges from the pattern across these five metrics.
- Detect-and-report rate: the share of employees who notice a social engineering attempt and report it to the organization. It is the core measure of awareness and reporting reflex. It shows the actual return of the awareness program.
- Time to report: the time from the start of the attempt to the first correct reporting action. It is the quantitative dimension of the reporting reflex. It is an input to the escalation percentage.
- Correct escalation percentage: the share of reports that reach the right channel, the right person, at the right time. It shows how the response flow actually works. It is an input to the process breakdown analysis.
- Process breakdown: verified breakdowns in the reporting flow, escalation chain, and response procedure. It is the evidence for the process leg of the prioritized work list.
- Control effectiveness: real-world effectiveness evidence for the email gateway, MFA, reporting channel button, and response platform controls. It is the backbone of the technical control recommendations.
Production steps
The five-metric evidence package is not born in a single wave. It is produced through a four-step discipline from design to collection.
- Campaign design: the target audience, Segment, wave plan, and scenario set are defined under written authorization. Channel selection and content approval are completed in this step.
- Control set definition: the process and technical control set to be verified (email gateway, reporting button, response flow) is defined as a baseline.
- Reporting channel verification: the employee's reporting channel (report-phish button, email, phone) is tested end to end. Channel capacity and the escalation chain are verified before measurement.
- Metric collection: the five metrics are collected broken down by Segment, persona, and wave. Anonymization is applied from the start. Individual names stay out of the report.
Reading example: reporting reflex in an anonymized campaign
In a two-wave targeted phishing campaign we ran in the high-volume services sector, the first wave produced a low detect-and-report rate, a long time to report, and a limited correct escalation percentage. In the process breakdown analysis, we confirmed that the reporting button fell into the help desk queue and did not flow directly to the response team. The control effectiveness finding pointed to process design. The training need landed in second place.
The numeric ranges are calibrated once client input is complete. The structure of the example narrative is preserved.
Decision impact: intervention points in the program
The five-metric measurement package produces its clear decision based not on which employee erred but on which point of the program broke down. The detect-and-report rate points to the training layer. Time to report points to reporting channel design. The correct escalation percentage points to the response flow. Process breakdown points to process design. Control effectiveness points to technical control configuration.
- If the detect-and-report rate is low: awareness program and micro-training design.
- If time to report is long: reporting channel accessibility and user experience.
- If the correct escalation percentage is low: response team routing and training.
- If there is a process breakdown: revision of the reporting, triage, and response flow.
- If control effectiveness is low: email gateway, MFA, and filter configuration.
Neighboring evidence objects
Human-layer program measurement and the attack-path human vector produce different evidence objects. The Social Engineering Simulation provides human-layer program evidence. The human-layer initial access package of the Modular Red Team Simulation, by contrast, works as a vector tied to the attack-path chain and is read in the context of progressing toward a critical target.
- Social Engineering Simulation: program measurement, a five-metric package.
- Modular Red Team Simulation human vector: attack-path chain evidence.
Cadence-based verification, planned re-verification, and risk trend: a picture of movement rather than a single report.
Continuous Penetration Testing goes beyond a one-time snapshot. It produces a picture of movement for leadership through monthly verification sprints, planned remediation verification, and quarterly trend. This card opens up the concrete components of that picture.
Evidence object definition
The evidence object is a triple backbone made up of monthly verification sprints, planned re-verification, a measurable remediation rate, and quarterly risk trend across a program lasting a minimum of three months. Beyond a single finding report, it is an evidence program carrying Cadence, remediation, and management visibility.
- Cadence calendar: the monthly verification sprint, the verification entitlement included in the package, and the distinction of out-of-Cadence release verification entitlement.
- Re-verification evidence: a controlled second verification after remediation, confirming the finding is closed via a request/response trace or PoC.
- Risk trend chart: the movement of new, remaining, and closed findings, plus a quarterly reading of MTTR and remediation rate.
- The remediate-and-verify loop: the monthly work list, remediation window, planned verification, and the next month's priority area.
- Dual-Layer delivery: a monthly summary and quarterly trend for leadership, and an owner-assigned work list for the technical team.
Production steps (monthly Cadence, remediation)
The evidence object is produced through a four-step loop repeated each month, aligned with the continuous verification dimension of the OWASP DevSecOps Maturity Model and the ASVS continuous verification approach. The output is a live program flow, not a static report.
- Monthly verification sprint: manual in-depth testing on the target asset and a scope sweep with a control set aligned to OWASP WSTG and the API Security Top 10.
- Remediation window: delivery of the monthly work list, a prioritized remediation label, and placement into the release plan owned by the product team.
- Planned re-verification: a controlled retest within the package entitlement, with remediation evidence (a request/response trace or the invalidation of a safe PoC).
- Trend update: a monthly executive summary and a quarterly trend chart. The next month's priority-area decision is clarified with leadership.
Reading example: a 6-month program trend
The structure below shows how the end-of-program trend reading over six months is derived. It contains no client, sector, or product name. Only the reading order carries the discipline of Cadence-based verification.
Target asset: web and API surface, two authenticated roles. Cadence monthly. Minimum 3 months. A total 6-month program.
New-finding volume is high in month 1, with a measured decline after month 4. The number of remaining critical findings gradually approaches zero.
The remediate-and-verify loop: the average remediation time for critical findings shortens gradually from month 2 to month 6. Every remediation carries a re-verification trace.
Quarterly trend: the drop in risk profile is visible. The next quarter's priority area was marked as the role expansion in the new release line.
Decision impact: management visibility
The evidence object is more than a single chart. The monthly executive summary and quarterly trend let leadership answer the "is risk going down?" question without reading countless lines. The output is an input to the next release decision and to budget prioritization.
- A picture of movement rather than a one-time snapshot. Leadership reads the answer to "how long has it been declining?"
- The monthly work list goes to the technical team as an owner-assigned work list. It touches the release plan.
- The quarterly trend report is in a portable format for the board or for a client security review.
Neighboring evidence objects
When the decision priority differs, the evidence object differs as well. The distinction below directs the reader to the right service. Continuous Penetration Testing is a distinct service, not the continuous version of another service.
- Application Security Penetration Testing: one-time deep verification, a prioritized remediation list, and a release, project, or review window.
- Modular Red Team Simulation: attack path, critical target impact, and detection gap matrix (a scenario-based engagement rather than a Cadence-based program).
Beyond a vulnerability list: a verified attack path, the chain to a critical target, the detection gap matrix, and the control recommendation that breaks the chain.
The output of the Modular Red Team Simulation goes beyond a finding list. What is produced is the distillation into an evidence file, together with a timeline, of which controls the attacker bypassed while advancing from an initial foothold to a critical target, which step the defense saw, and which step it missed.
Evidence object definition: what the attack-path package covers
The evidence object consists of three separate layers, delivered bound to the same scenario: the attack-path timeline, the detection gap matrix, and the chain-breaking control recommendation. Rather than three separate files, they are three faces of the same scenario.
Attack-Path Timeline: the steps advancing from the initial foothold (threat actor access, critical asset targeting, or an assumed breach start) to the critical target are recorded together with a timestamp, the technique used (MITRE ATT&CK technique ID), and the privilege level obtained. The chain can be read in the order of "start, reconnaissance, initial access, privilege escalation, lateral movement, persistence (controlled), and target impact."
Detection Gap Matrix: each step in the attack path is tagged with a MITRE ATT&CK technique ID. Whether the step was detected in SOC and EDR/SIEM telemetry, which alert level it fell into, and whether it reached correct escalation are recorded row by row. The gap is read in three types: no record, record but no alert, alert but escalation missing.
Chain-Breaking Control Recommendation: beyond controls that close the attack path one by one, a control set that breaks the chain at the most economical point is recommended. The recommendations are mapped to the MITRE D3FEND defensive technique catalog and enter investment prioritization with the clarity of "if these three steps are closed, the chain breaks."
How the evidence is produced
The production flow runs in four phases aligned with the TIBER-EU and CBEST frameworks: the actor profile is derived through threat intelligence, scope and target are fixed through scenario design, the attack path is executed through controlled execution, and findings are handed over to the defense team through an assessment and purple team session. The NIST SP 800-115 manual verification discipline is preserved at every step.
- Threat intelligence: an actor profile fit to the organization, sector threat landscape data, and targetability analysis.
- Scenario design: core module selection (RT-1, RT-2, RT-3), crown jewel definition, and the starting assumption.
- Controlled execution: written authorization, narrow stakeholder briefing, approval gates, and activity logging.
- Assessment and purple session: handover of findings to the defense team, detection rule mapping, and a remediation verification plan.
Reading example: anonymized case
A target-focused scenario at a high-volume services-sector organization. After defining the corporate payment reconciliation system as a critical asset, we documented, with a seven-step timeline and MITRE ATT&CK technique mapping, the chain that started from a limited identity obtained through an external-facing developer portal and reached the target system by way of an old VPN privilege into the internal network and then a misconfigured service account ticket. The SOC alerted on two of the steps, three steps had a record but no alert, and two steps were entirely off the record.
This reading example is structural. The sector category, product name, technology name, and geographic identifiers have been anonymized. The full evidence file template is expanded on the Evidence Package Examples page.
Decision impact: how the attack path opens up a decision
When the evidence object reaches the investment table, three decision tracks light up together: which detection gap warrants telemetry investment, which control set breaks the attack path at the most economical point, and what the top three items of the resilience work list will be. Dual-Layer delivery carries the same evidence to the senior leadership risk table and the technical team work list with the same clarity.
- Detection investment priority: for which technique will the record or alert gap be closed?
- Chain-breaking control: which three controls, when closed, break the attack path most economically?
- Resilience work list: which steps tie into the remediation verification loop, and which into re-verification?
Neighboring evidence objects: evidence that looks close but stands apart
In Network Security Penetration Testing the evidence object is a network surface finding: a verified technical flaw and segmentation impact across the external, internal, and wireless layers. In the Modular Red Team Simulation the evidence is the attack path: the chain to a critical target and the detection gap. The two evidence objects are related but not synonymous. A network surface finding can feed the chain but carries no promise of an attack path.
The evidence object of Generative AI Red Team is the GenAI runtime chain: the abuse path of agent behavior across the prompt, tool, data, and privilege layers. While the Modular Red Team Simulation targets the organization-wide attack path, the GenAI Red Team verifies the runtime chain of a productized agent or a RAG/MCP integration. The two services are not discussed on the same page.
A proven abuse path across the prompt, tool, RAG, MCP, privilege, and action chain, or the control point that breaks the chain.
Generative AI Red Team produces an output that goes beyond a one-line prompt injection note: reproducible evidence and a chain-breaking control recommendation across the live sequence of prompt, tool call, RAG retrieval, MCP authorization, and privilege decision within the agent workflow. This card opens up the concrete components of that package.
Evidence object definition
The evidence object is a behavior finding demonstrated reproducibly across the live chain of the agreed agent workflow, running from prompt to tool, from tool to data source, and from there to the privilege and approval decision, together with visibility into the affected data and action class, plus a chain-breaking control recommendation ordered by business impact. Beyond a single jailbreak list, it is a decision-ready evidence file.
- Reproducible evidence: a safe PoC, prompt/response trace, tool call log, and MCP token flow dump.
- Impact map: a clear explanation framed around agent privilege, RAG disclosure, tool abuse, and MCP authorization exploitation.
- OWASP Top 10 for LLM Applications 2025 and MITRE ATLAS mapping: which risk family is in play at which link of the chain.
- Dual-Layer delivery: a risk picture for leadership and an owner-assigned work list for product and engineering.
Production steps (MITRE ATLAS and NIST AI RMF aligned)
The evidence object is produced in a flow aligned with the MITRE ATLAS tactic and technique order, the NIST AI Risk Management Framework GOVERN/MAP/MEASURE/MANAGE functions, and the OWASP Top 10 for LLM Applications 2025 risk families. The output is a behavior verification pipeline, not merely a document.
- Scope and rules of engagement verification: agent workflow, tool catalog, RAG sources, MCP servers, and the role and approval boundary.
- Attack surface derivation and threat modeling (NIST AI RMF MAP and MITRE ATLAS Reconnaissance / Initial Access).
- Prompt, RAG, tool, and MCP testing: indirect prompt injection, sensitive information disclosure, tool abuse, and MCP authorization exploitation.
- Finding confirmation and risk prioritization (NIST AI RMF MEASURE; business impact and the link of the chain are evaluated together).
- Executive summary, technical report, PoC production, and chain-breaking control recommendations (NIST AI RMF MANAGE).
- Optional remediation verification and retesting (focused or end-to-end retesting).
Reading example: anonymized finding structure
The structure below shows how a single finding record from a Generative AI Red Team report is read. It contains no client, sector, or product name. Only the reading order carries the discipline of the evidence object.
Agent privilege leakage via indirect prompt injection: unauthorized invocation of a write-capable tool through a hidden directive in a RAG source.
The agent invoked the record-update tool without user approval. The tenant boundary was violated and a sensitive field was modified.
Prompt/response trace, tool call log, and safe PoC. Repeated across three separate sessions, with the MCP token flow recorded.
OWASP LLM01 Prompt Injection, LLM06 Excessive Agency, and LLM08 Vector and Embedding Weaknesses; MITRE ATLAS Initial Access and Execution.
Critical severity. The first item of the pre-release chain-breaking control, with an approval gate mandatory for the write-capable tool.
Decision impact: the chain-breaking control
The output is more than a finding list. Each item comes with a control recommendation specifying which link of the chain it sits on. This lets the product team make the decision to harden the system prompt, tool permissions, RAG access, MCP approval, or output filter based on data.
- System prompt hardening: safe execution boundaries, instruction hierarchy, and untrusted flagging for indirect content.
- Tool permission narrowing: an approval gate for write-capable tools, least privilege, rate limiting, and audit logging.
- RAG access control: source verification, content filtering, per-user access control, and a disclosure filter.
- MCP approval gate and token management: authorization scope checks, session limits, and blocking token forwarding.
- Output filter and observability: sensitive data redaction, anomalous tool call detection, and log streaming.
Neighboring evidence objects
When the decision priority differs, the evidence object differs as well. The distinction below directs the reader to the right service. The three-way AI distinction does not melt under a single roof.
- AI Model Supply Chain Assurance: source provenance, component list, attestation, and registry trust.
- AI for Penetration Testing and Red Team: internal offensive team workflow design. Not an external client test.
The evidence object is a five-component model supply trust package: source provenance, BOM, attestation, registry signature, and release pipeline trust.
The primary evidence object of AI Model Supply Chain Assurance is not reduced to a single signature. Five components are read together, and the defensible trust ground for the release decision emerges from the pattern across these five components.
Evidence object: the five-component package
- Source provenance: a verifiable record of where the model artifact, dataset, code, dependencies, and build outputs came from. Beyond just the "where did the model come from?" question. It is the provenance of the data and dependencies too.
- Bill of materials (BOM / AI-ML-BOM): a structured component list for the model artifact, data source, dependencies, and build outputs. The CycloneDX and SPDX schemas, with the AI-ML-BOM extension, cover model-specific components (weights, tokenizer, training data reference).
- Attestation: signing the outputs of the build/CI pipeline and recording attestations in a structured way. It rests on the in-toto, Sigstore, and Cosign foundation. The promote/rollback decision is tied to these attestations.
- Registry signature: keeping model artifacts and versions signed in the central registry, with access control, promote/rollback, and audit log evaluated together. Abuse scenarios (unauthorized promote, silent rollback) are read against this evidence.
- Release pipeline trust: the trust chain of the path by which the model is taken from the registry into the deployment environment. Across staging and production, in multi-tenant and multi-region topologies, the evidence chain is verified for every transition. The manifest signature and verification threshold are the control point.
Production steps
The five-component evidence package is not born in a single wave. It is produced through a six-step discipline from data source to deployment manifest.
- Data source chain: the source, license, and integrity of training and fine-tune data are recorded. A BOM entry is opened for third-party datasets. The data provenance chain is verified as the first link.
- Training/fine-tune chain: the training pipeline, hyperparameter configuration, base model version, and fine-tune steps are tied to attestation. The privilege boundary of the build/CI pipeline and the attestation input are clarified in this step.
- Artifact and dependency BOM: a CycloneDX- or SPDX-formatted BOM is produced for the model artifact (weights, config, tokenizer, template), code, and library dependencies. Lockfile and secret scanning outputs are added to the same evidence package.
- Registry signing and audit log: when the artifact is uploaded to the registry, Sigstore- or Cosign-style signing is applied, and promote/rollback actions are recorded in the audit log. Access control and the role matrix are proven in this step.
- Deployment manifest verification: the model version, signature chain, and attestation reference are verified in the release manifest. An artifact that does not pass the verification threshold is not promoted to production. The environment-based evidence chain is preserved.
- Evidence package consolidation: the five-component evidence chain is consolidated into an executive summary, a BOM appendix, and an attestation gap analysis. It is delivered as a portable package for client review and audit.
Reading example: evidence chain gap of an anonymized fine-tuned model
In an assessment we ran on a release-ready fine-tuned model at a high-volume sector platform, three components came up missing: there was no training data BOM record, no attestation pipeline had been set up for the fine-tune step performed on the base model, and the registry signature covered only the model artifact, leaving the training data and dependency layer out. Because the signature chain in the release manifest rested only on the final artifact, source provenance could not be verified. We identified the evidence gap before the release decision. The promotion to production was withdrawn as a precaution.
The numeric ranges and sector attributes are calibrated once client input is complete. The structural integrity of the example narrative is preserved.
Decision impact: the trust basis of the release decision
The five-component model supply trust package produces its clear decision by answering not "does the model work?" but "on what basis are we releasing the model?" A source provenance gap points to the data supply pipeline. A BOM gap points to dependency hygiene. An attestation gap points to the build/CI pipeline. A registry signature gap points to the access and promote/rollback track. A release pipeline weakness points to deployment manifest verification.
- If source provenance is missing: the data supply chain and license inventory are opened.
- If the BOM is missing: dependency hygiene and a CycloneDX/SPDX pipeline are set up.
- If attestation carries a gap: build/CI signing and an in-toto/Sigstore pipeline are brought in.
- If the registry signature is missing: access control and promote/rollback hardening are performed.
- If the release pipeline is weak: a deployment manifest verification threshold is defined.
Neighboring evidence objects
Model supply trust and the runtime chain produce different evidence objects, while the internal offensive workflow is an entirely different class of evidence object. AI Model Supply Chain Assurance proves the provenance, integrity, and trust of the model and the release pipeline. Generative AI Red Team provides runtime-chain evidence across prompt, tool, data, and privilege. AI for Penetration Testing and Red Team names the improved internal delivery workflow.
- AI Model Supply Chain Assurance: the model supply trust package.
- Generative AI Red Team: runtime chain evidence.
- AI for Penetration Testing and Red Team: internal workflow design.
The evidence object is a five-component internal delivery package: workflow design, approval model, eval set, observability, and controlled pilot.
This service does not produce the result of an external attack. Within the internal offensive delivery process, the workflow design, approval model, eval set, observability, and controlled pilot output package are delivered together as evidence.
Evidence object: the five-component internal delivery package
The primary evidence object of AI for Penetration Testing and Red Team is not reduced to a single acceleration figure. Five components are delivered together, and the clear decision of the internal pilot emerges from the pattern across these five.
- Workflow design document: the master document that structures the AI-assisted form of a use case in the delivery process (reconnaissance, hypothesis, evidence processing, reporting, or retesting) at the granularity of a use case unit, with a clear owner and acceptance criteria.
- Approval model: a written map showing who approves each step of the workflow, at what privilege level, and with what acceptance logic, covering the use case, RoE, artifact/access, action level, PoC acceptance, and pilot acceptance gates.
- Evaluation set (eval set): a structured set of test scenarios, verifiable metrics, and acceptance gates designed to test the AI-assisted workflow against acceptance criteria for accuracy, security, and delivery quality.
- Observability design: recording each step of the workflow with logging, audit trail, telemetry, and rollback criteria, making the engagement open to human oversight and Evidence-Based traceable.
- Controlled pilot output package: the recorded output of running the designed workflow through owner-assigned sprints within the approval model, eval set, and observability. It is the delivery package of the third package level that establishes a reproducible workflow standard.
Production steps
The five-component evidence package is not born in a single wave. It is produced through a five-step discipline from use case selection to handover.
- Use case selection: a single workflow problem in the offensive delivery process is bounded to one main output, one approval model, and one owning team. Instead of an abstract "leveraging AI," we drill down to a concrete use case unit.
- Fixing the approval model: the approval owner, privilege level, and acceptance logic are put in writing for each stage of the workflow. Uncontrolled privileged action and autonomous production action are rejected. Human decision is preserved at every gate.
- Setting up the eval set: the eval set is built with accuracy, security, and delivery quality metrics. Acceptance ranges, stop thresholds, and quality-loss signals are written into the eval set from the start.
- Controlled pilot: the workflow is run as a controlled pilot within a sandbox or canary engagement. The observability design records every step. A drop in the eval set triggers the pilot stop rule.
- Handover: when the pilot is accepted, the workflow document, approval model, eval set, observability design, and operations guide are handed over to the internal team. The next sprint plan is written together.
Reading example: speed versus quality in an anonymized internal pilot
On an internal offensive team that regularly delivers penetration tests, we ran an evidence-to-finding workflow use case pilot. The AI-assisted workflow normalized screenshots, logs, and PoC evidence and turned them into a finding draft. The expert analyst verified the final finding decision every time. When the eval set fell below the accuracy threshold, we stopped the pilot, revised the draft, and restarted. The acceleration was qualitative. Expert decision authority was not moved outside the process.
The numeric ranges are calibrated once the internal pilot is complete. The structure of the example narrative is preserved.
Decision impact: moving expertise into the system
The five-component evidence package produces its clear decision by answering not "how much did AI accelerate things?" but "which step was moved into the system while preserving expert quality?" The workflow design brings reproducibility. The approval model closes the privileged-action risk. The eval set makes the quality threshold visible in numbers. Observability provides auditability. The controlled pilot delivers the operational evidence of this whole structure.
- The workflow is reproducible. Person-dependent delivery decreases.
- The approval gates are written. The uncontrolled-action risk closes.
- The eval set catches quality loss early, in numbers.
- Observability establishes the ground for auditing and rollback.
- The controlled pilot turns the workflow standard into written evidence.
Neighboring evidence objects
Three different evidence objects must not be confused under the AI heading. AI for Penetration Testing and Red Team names the internal offensive workflow. Generative AI Red Team tests the runtime chain of the client's product. AI Model Supply Chain Assurance assesses the model's source provenance and release pipeline trust. The evidence object of the three tracks differs. It cannot be flattened under a single heading.
- AI for Penetration Testing and Red Team: internal workflow evidence.
- Generative AI Red Team: runtime chain evidence.
- AI Model Supply Chain Assurance: model supply trust evidence.
Controlled testing engagement and preparation
Controlled testing frameworks
Each service's authorization, scope, RoE, approval gates, and escalation discipline.
View the frameworks PreparationPreparation and scoping guides
The prerequisites, access, and stakeholders you need to prepare before testing begins.
View the guides ExampleAnonymized evidence package examples
The anatomy of an evidence package: an anonymized example for three services.
View the examples// EVIDENCE OBJECT
Which evidence object closes your decision?
Share your decision priority, and let us clarify the right service and the first evidence object together in a discovery call.