SOC 2 Compliance for Engineering Teams: What Actually Matters

Most SOC 2 prep focuses on policy theater. Auditors care about code-level controls: PR reviews, secrets management, deployment gates, and audit trails that prove your access controls actually work.

Connectory team|Updated 17 min

Picture this: your company just spent six months preparing for SOC 2. You hired a consultant, wrote 47 policies, and conducted weekly readiness meetings. Then the auditor asks to see evidence of your PR approval enforcement from last quarter. Your team scrambles through GitHub, Slack, and Jira trying to reconstruct who approved what and when. Three weeks later, you're still collecting screenshots.

This is the cost nobody warns you about. Consulting and audit fees are only part of it; the hidden cost is the engineering time spent proving your controls actually work. The real work isn't writing policies, it's instrumenting your SDLC to produce compliant evidence automatically, every single day.

Teams can struggle in Type II audits not because their security is weak, but because they treated SOC 2 as a documentation exercise instead of an engineering challenge. The controls worked. They just couldn't prove it.

The Documentation Tax Nobody Warns You About

SOC 2 auditors don't just read your access control policy and move on. They sample pull requests, deployments, and access logs to verify that your documented procedures match reality. If your policy says "all production changes require two approvals," they can pull PRs from March, July, and November to confirm each one had two reviewers. A self-approval or force push becomes an exception in the report.

Many audit problems are documentation gaps, not missing technical controls. Your infrastructure is secure. Your secrets are vaulted. Your deployments are gated. But when the auditor asks for quarterly access reviews from Q2, you realize nobody exported the data. The control existed. The evidence didn't.

Teams that treat SOC 2 as policy theater struggle when auditors start sampling. A first audit takes far more engineering time when the SDLC was built without evidence generation in mind. Every PR approval, every deployment, every secrets rotation needs a timestamped, immutable record linking the action to a specific human identity. Manual screenshots don't scale to sampling across a full observation period.

The engineering teams best placed for Type II audits are the ones who built evidence collection into their normal workflow from day one. Compliance becomes a side effect of proper instrumentation, not a separate manual process that derails sprints every quarter.

Your CI/CD pipeline already captures most of what auditors need, deployment timestamps, test results, approval gates. The gap is extracting that data in auditor-friendly formats without burning hours per control every quarter. This is an engineering problem with an engineering solution, but most teams don't realize it until they're three weeks into audit response.

The Five Controls Auditors Actually Test in Your Codebase

SOC 2's Trust Services Criteria cover much more than engineering, but five areas map most directly to your engineering workflow. These aren't abstract policy requirements, they're testable controls with specific evidence artifacts.

CC6.1 (Logical and Physical Access Controls) is where engineering teams usually point to PR approval enforcement, branch protection rules, and 2FA on GitHub or GitLab. If those are your documented controls, expect the evidence to be checked: that your main branch requires reviews, that force pushes are disabled, and that developer accounts have multi-factor authentication enabled. A developer account without 2FA is the kind of gap that can turn into an exception. How your auditor tests each control depends on the controls you describe in your system description.

Secrets management sits under the logical access criteria. Expect your commit history to be scanned for hardcoded credentials with tools like TruffleHog and GitLeaks. A single API key committed to a private repo in 2023 and removed in 2024 can still be a finding. Expect questions about secrets rotation schedules, access logs showing who read production credentials, and evidence that secrets live in HashiCorp Vault or AWS Secrets Manager, not environment files in your repo.

SOC 2-Compliant Deployment Pipeline with Evidence Generation
SOC 2-Compliant Deployment Pipeline with Evidence Generation

CC7.2 (System Operations) covers monitoring and incident response. Auditors want deployment logs, error tracking timestamps, and evidence that you actually respond to security alerts. If your policy says "critical alerts escalate within 15 minutes," they'll sample PagerDuty logs to confirm escalation timing. Your monitoring exists, but can you prove response times?

CC8.1 (Change Management) is the control most teams underestimate. If your documented change process routes every production change through a reviewed pull request, the evidence has to show that it did. Console cowboy commits, SSHing into a server and editing config files directly, break that rule. Auditors will cross-reference deployment logs against your PR history to find changes that bypassed review. An emergency hotfix pushed without a PR shows up as a change management exception.

Availability criteria apply when availability is in scope for your report, and they cover your deployment gates and rollback procedures. Expect requests for evidence that failed deployments roll back, that health checks prevent bad releases from reaching users, and that your runbooks actually work. Rollback procedures need to be tested, not theoretical.

The One Question That Determines Audit Readiness
Can you produce evidence for any control, for any random date in the last 12 months, in under 30 minutes? If not, you're not ready for Type II. Manual evidence collection doesn't scale to quarterly sampling across hundreds of controls. Build automated evidence pipelines before the auditor shows up.

Why AI-Generated Code Creates New Compliance Gaps

SOC 2 frameworks were written before AI coding assistants existed. Auditors are starting to ask questions the Trust Services Criteria don't address directly: Who approved the AI to touch this production endpoint? How do you review code when the original author is a language model?

In CodeRabbit's analysis of 470 open-source pull requests, AI-co-authored PRs produced 10.83 issues per PR against 6.45 for human-only PRs, about 1.7x more [1]. The code looks correct but contains subtle bugs that slip past cursory review. For SOC 2 purposes, this means your PR review control needs explicit evidence that human eyes examined AI output. A rubber-stamp approval on Copilot-generated code isn't sufficient if the reviewer spent 12 seconds on a 400-line PR.

The deeper problem is identity. CC6.1 requires access controls tied to specific individuals. AI tools add identities and actions that centralized IAM may not see. When Cursor autonomously refactors your authentication module in agent mode, which human authorized that change? Traditional audit trails show "Developer A merged PR #847," but they don't capture "Developer A's AI agent autonomously modified 14 files while Developer A was in a meeting."

Non-human identities are creating a compliance gap. Your IAM system knows Developer A has production access. It doesn't know Developer A's Copilot session can read production secrets, or that their Claude Code instance has write access to infrastructure-as-code repos. The credentials exist. The audit trail doesn't.

In practice, this means AI coding tools need to be treated as non-human identities with explicit access grants and audit logging. When a developer enables Cursor's agent mode on a codebase containing PII, that authorization decision should be logged. When Copilot suggests code that calls a payment API, the approval of that suggestion should tie back to a human reviewer with documented authority to modify payment workflows.

Teams preparing for SOC 2 can instrument their AI tool usage: log assistant sessions where tools support it, and keep audit trails of which files an agent modified. The PR review control then needs human approval plus evidence that the reviewer actually examined the AI-generated portions instead of assuming correctness.

The PR Review Control That Passes (or Fails) Your Audit

Pull request review is one of the most visible engineering controls in a SOC 2 audit. Auditors sample PRs across your observation period to verify that mandatory approval enforcement worked consistently. Each sampled PR that missed it is an exception.

Branch protection rules must match your documented policy exactly. If your policy says "two approvals required for production code," GitHub's branch protection settings must enforce two reviews. But here's where teams fail: your policy says "two independent reviewers," but GitHub allows any two people to approve, including the PR author's alternate account or a junior developer rubberstamping their manager's code.

Self-approvals, force pushes that bypass review, and admin overrides that merge without approval all produce exceptions. Your GitHub audit log will show every bypass, and auditors know where to look.

The evidence artifact auditors want is a timestamped record showing reviewer identity, approval timestamp, and merge authorization. This lives in GitHub's API, but most teams don't export it until audit season. Then they discover PRs from Q2 where the approval came from a user account that no longer exists, or a bot that shouldn't have had approval authority.

Automated code review adds a compliance layer that helps with audit evidence. A tool like SlopBuster performs machine pre-review, flags issues, and generates a review comment before human approval. This creates a two-tier audit trail: automated analysis ran and found X issues, then Human Reviewer Y approved after those issues were addressed. The timestamp gap between machine review and human approval proves the reviewer had time to actually examine findings.

PR Review ConfigurationAudit RiskEvidence QualityRemediation Effort
No branch protectionCriticalNone - all evidence manualLong reconstruction of approvals
Branch protection without required reviewsHigh - shows intent but no enforcementPartial - need to manually verify each PRSignificant manual sampling
Required reviews but allow self-approvalHigh - policy/practice mismatchGood logs, bad practicesWork to explain historical violations
Required reviews + automated scanningLow - dual verification layerExcellent - machine + human trailExporting evidence
Full enforcement + review time trackingMinimalExcellent + shows review timeAutomated export

The detail that catches teams off-guard: auditors want proof that reviewers spent meaningful time examining code, not just clicking "approve." If your GitHub data shows PRs approved 8 seconds after creation, auditors will question whether real review occurred. This is where PR complexity metrics matter, large, complex PRs need longer review times to pass the "reasonableness" test.

Teams using AI-assisted review need to be especially careful. If your process is "Copilot writes code → developer approves without reading → teammate rubberstamps," you're creating a paper trail of negligent review. Better approach: AI writes code → automated scanner flags concerns → human reviews AI output and scan results → second human approves after issues addressed.

Secrets Management: The Control Most Teams Fail

Secrets management violations are the easiest audit findings to detect and the hardest to remediate. Hardcoded credentials anywhere in your commit history can become findings, even if you removed them three years ago. Git never forgets.

Scanners such as GitLeaks and TruffleHog find API keys, database passwords, AWS access tokens, and private keys in commits dating back to your first push. The remediation isn't "remove the secret", that file still exists in Git history. OWASP notes that squashing history to remove an exposure "may introduce other problems as it rewrites git history" [2]. Proper remediation requires rotating the compromised credential and proving the rotated secret is stored securely.

OWASP's guidance is direct: "You should regularly rotate secrets so that any stolen credentials will only work for a short time" [2]. Your rotation logs should show secrets changed on the schedule your policy sets, with access restricted to named individuals. This is where automation pays off, because manual rotation drifts and is hard to prove. Vault's audit devices, for example, "record all API requests and responses in detail" with a small set of exceptions [3].

Environment variable management creates a hidden compliance gap. Your production secrets live in Vault, but how do they reach your application? If developers can echo $DATABASE_PASSWORD in a production shell session, that undercuts the access controls you document. Your evidence needs to show that only the application runtime can read secrets, not the humans who deployed it.

CI/CD secrets injection must be auditable. Many teams have GitHub Actions workflows that access production credentials via repository secrets. Who can read those secrets? GitHub's audit log knows, but you need to export it. Who can modify workflow files to exfiltrate secrets? Your branch protection controls that. One developer with write access to .github/workflows but no documented authorization to read production secrets is an access control gap.

1.7x
More issues per PR in AI-co-authored changes than in human-only PRs, per CodeRabbit [1]

A scenario that creates findings: Your team uses AWS Secrets Manager for production databases. A developer needs to debug a production issue, so they aws secretsmanager get-secret-value from their laptop. The secret was accessed by an authorized individual for a legitimate purpose, but there's no ticket, no approval, and no automated revocation when the debug session ended. The audit trail shows secret access. It doesn't show authorization, purpose, or time-limited grant.

Better pattern: Developers never read production secrets directly. Production debugging uses session management tools (AWS Systems Manager Session Manager, Teleport) that audit every command. If a developer needs to verify a database connection, they get a time-limited credential that expires in 1 hour and generates an audit log entry. The secret itself never touches their laptop.

Building Audit-Ready Deployment Gates Without Killing Velocity

The myth is that SOC 2 compliance requires slow, bureaucratic deployment processes. When controls are automated in the pipeline, compliance and velocity aren't opposites.

Every production deployment needs five pieces of evidence: who initiated it, what changed, when it deployed, why it was approved, and proof of authorization. This evidence should be automatically generated, not manually documented. Your CI/CD pipeline already captures most of this, the challenge is structuring it for auditor consumption.

Deployment gates should be automated checks plus human approval. Automated checks: tests pass, security scan clear, no high-severity vulnerabilities, deployment to staging succeeded. Human approval: authorized individual reviewed the change and confirmed deployment timing. The automated checks create objective evidence. The human approval creates accountability.

Rollback procedures need documentation and test evidence. Your runbook says "redeploy previous version via CI/CD rollback job," but when did you last test that procedure? Auditors will ask for proof that rollbacks work. Some teams schedule quarterly rollback drills and save the Jenkins logs as compliance evidence. The drill tests the procedure, the logs prove it was tested.

GitOps patterns naturally create compliant audit trails as a side effect of infrastructure-as-code. Every environment change is a Git commit. Every commit has an author, timestamp, and review history. Your Kubernetes manifests in Git are your deployment documentation. ArgoCD or FluxCD sync logs are your deployment evidence. The infrastructure enforces the control automatically.

The detail auditors probe: whether deployment automation actually prevents human bypass. If your process requires a PR review but developers can still kubectl apply directly to production, the control is documented but not enforced. Auditors test for this by checking cluster RBAC permissions. Who has direct production access? Why? Is it documented and reviewed quarterly?

Teams that maintain velocity under SOC 2 treat compliance as code. Their deployment gates are automated policy checks (OPA, Kyverno, Sentinel) that enforce requirements without requiring manual approvals for every change. Their evidence collection is a background process that exports relevant data to a compliance dashboard. Their quarterly access reviews are database queries, not spreadsheet archaeology.

The Three Documentation Artifacts Auditors Demand to See

Three evidence types determine audit outcomes. Miss any of them, and you're scrambling during the audit to reconstruct history.

Access review logs prove you quarterly reviewed who has production access and why. This isn't a spreadsheet showing current permissions, it's a timestamped record showing you examined permissions on March 15, June 15, September 15, and December 15, made decisions about appropriateness, and revoked unnecessary access. The evidence includes the reviewer's identity, the date of review, and actions taken.

Most teams fail this control because they review access informally. The engineering manager knows who should have production access, but there's no record of the quarterly check. When auditors ask for Q2 access review evidence, the team realizes they reviewed access in a Slack thread that's been deleted.

Better approach: Access reviews are Jira tickets with a checklist of every user and service account with production permissions. The ticket includes approval from security or management, a deadline for completion, and a comment thread showing decisions. The closed ticket is your evidence.

Incident response records must include timestamps, communication logs, and resolution evidence for every security event. "We had a security incident in July and fixed it" isn't sufficient. Auditors want: detection timestamp, initial responder, escalation times, communication with affected parties, root cause, remediation steps, and verification that the fix worked.

Your PagerDuty, Slack, Jira, and post-mortem documents combined create this evidence, but only if you collect them at the time. Trying to reconstruct a security incident six months later from memory and scattered Slack threads rarely produces evidence an auditor accepts. Teams that pass create an incident response template and fill it out during the response, not after.

Change management tickets tie every production change to an approved work item with reviewer identity. This is the control where AI code creates new challenges. If your GitHub PR was authored by Copilot, who approved the AI to make that change? The PR reviewer approved the code, but did they authorize the AI to touch that part of the system?

In practice, change management evidence is your PR history plus your project management system. Jira ticket ABC-123 describes the work. GitHub PR #456 implements it. The PR references the ticket. The ticket links to the PR. The deployment references both. This creates a bidirectional audit trail from business requirement to deployed code.

Most failures happen when engineers can't produce evidence for controls they claim to follow. The access reviews happened, but there's no record. The incident response was effective, but the documentation was verbal. The changes were approved, but the approval was a hallway conversation, not a tracked decision.

Automated evidence collection transforms this. Instead of manually exporting GitHub PRs every quarter, set up a weekly job that extracts PR metadata, approval timestamps, and reviewer identities to a compliance database. Instead of Slack threads about access reviews, create a scheduled workflow that generates a review ticket, pre-populates the user list, and requires sign-off. Instead of post-incident documentation archaeology, use a PagerDuty or Opsgenie integration that auto-creates the evidence artifact when an incident closes.

Evidence Collection MethodTime Per QuarterReliabilityEvidence ConsistencyBest For
Manual screenshotsHeaviestLow - human errorMedium - formatting inconsistenciesStartups pre-SOC 2 with <10 engineers
Quarterly exports from toolsHeavyMedium - depends on memoryGood - structured dataGrowing teams preparing for first audit
Automated weekly collectionLight reviewHigh - runs automaticallyExcellent - consistent formatType I ready teams planning Type II
Continuous compliance pipelineMinimal validationVery high - real-timeExcellent - queryable on demandType II and beyond

Preparing for Type II: What Changes Between First and Second Audit

Type I SOC 2 is a point-in-time snapshot. Type II tests continuous control operation over an observation period that usually spans months. The difference isn't just duration, it's the sampling methodology that catches teams off-guard.

Type I auditors verify your controls work right now. They check that branch protection is enabled, test a few PRs, confirm your secrets are vaulted. Type II auditors sample controls across the entire observation period. They'll pull random PRs from March, June, September, and December to confirm that two-reviewer requirement was enforced consistently. One missing approval in a July PR becomes an exception in the report.

This is why automated evidence collection becomes mandatory for Type II. Manual screenshots don't scale to quarterly sampling. When auditors ask for evidence that all production deployments in Q3 followed the documented approval process, you need a database query that exports the data in 5 minutes, not a three-week scramble through Jenkins logs.

Integration between tools creates the compliant data pipeline. GitHub webhooks send PR events to a compliance database. HashiCorp Vault audit logs stream to your SIEM. PagerDuty incident events trigger evidence collection workflows. AWS CloudTrail feeds into a data warehouse that aggregates access patterns. This isn't a separate compliance system, it's engineering intelligence infrastructure that happens to produce SOC 2 evidence as a side effect.

The teams that struggle with Type II are the ones who treated Type I as a one-time documentation sprint. They wrote the policies, configured the tools, and passed the point-in-time check. But they didn't build continuous evidence collection, so when Type II sampling starts, they're reconstructing history instead of querying a database.

The teams that pass Type II easily are the ones who built compliance observability from day one. Their Engineering Intelligence Dashboard shows real-time metrics on PR review compliance, secrets management hygiene, and deployment gate effectiveness. They spot control drift immediately, if PR approvals start taking too long or developers begin bypassing reviews, alerts fire before it becomes an audit finding.

Quarterly preparation time for Type II depends on your control scope and evidence systems. Export the quarterly sample data, validate the sampled controls, investigate anomalies, and generate the auditor-ready report. If quarterly prep takes multiple person-weeks, identify which evidence still requires manual reconstruction.

The shift from Type I to Type II is the shift from "prove your controls exist" to "prove your controls work consistently across time." The latter requires engineering rigor, not policy documentation. Build observability, automate evidence collection, and treat compliance as a continuous process. The audit becomes a formality instead of a crisis.

---

Automate compliance evidence with code governance. SlopBuster provides SOC 2-ready audit trails for every pull request, with automated security scanning and merge gate enforcement. Learn how our compliance solution helps regulated teams, explore security features, or see how CISOs use Connectory for AI code governance at scale.

References

[1] CodeRabbit, "State of AI vs. Human Code Generation Report," 2025. https://www.coderabbit.ai/blog/state-of-ai-vs-human-code-generation-report

[2] OWASP, "Secrets Management Cheat Sheet." https://cheatsheetseries.owasp.org/cheatsheets/Secrets_Management_Cheat_Sheet.html

[3] HashiCorp, "Vault audit devices." https://developer.hashicorp.com/vault/docs/audit