Test Pyramid
A broad mocked unit-level suite (pytest, every GitHub API call mocked via respx), organized one
file per source module (tests/test_<module>.py), plus two cross-cutting files:
tests/test_idempotency.py (one integration-shaped test proving a second apply against an
already-compliant repo makes zero mutating calls) and tests/test_policies_parity.py (a
parametrized regression guard: for every field in diff._FIELDS, both the branch_protection and
ruleset backends must respond to it). Run pytest --collect-only -q | tail -1 for the exact,
current count — it changes with every test added, so it’s deliberately not hardcoded here (see
“__version__/PyPI version divergence” below for what hardcoding a value that’s supposed to track
something else costs this project).
Alongside the mocked suite, 3 end-to-end test functions in tests/e2e/ exercise the real GitHub
API against a live, persistent fixture repo (shipsolid/repo-policy-e2e-fixture): two independent
smoke tests (config validation, token/repo reachability) and one collapsed lifecycle scenario
covering everything else (see “E2E Test Isolation and Concurrency Safety” below for why it’s one
test, not several). The E2E suite is excluded from the default pytest run (pytest marker e2e);
run it explicitly with pytest -m e2e (requires REPO_POLICY_E2E_TOKEN), or via the
nightly/manual .github/workflows/e2e.yml.
Approach: TDD throughout
Every module in src/repo_policy/ was built test-first: write the failing test, watch it fail for
the expected reason, implement the minimal code to pass, run the full suite, commit. The TDD
discipline itself is enforced by convention, not tooling; the resulting coverage level is enforced
by CI (ci.yml’s test job fails under 95% branch coverage — see Known Gaps for what CI does and
does not gate).
What the mocked suite is good at
- Every GitHub API interaction is mocked via
respxagainst real-shaped fixture payloads (built from GitHub’s actual documented request/response schemas, not guesses) — seetests/test_github_client.pyfor retry/pagination/auth coverage. - The diff/resolve engine (
tests/test_diff.py) and both backend translators (tests/test_policies_branch_protection.py,tests/test_policies_rulesets.py) are exercised field-by-field, including the polarity inversion ofallow_force_push/allow_deletionrelative to every other field. tests/test_idempotency.pyproves, at the unit level, that applying an already-compliant policy twice makes zero mutating calls the second time — the tool’s core correctness promise.tests/test_policies_parity.pyis a parametrized guard: for every field indiff._FIELDS, both thebranch_protectionandrulesetbackends must respond to it. This exists specifically because a bug once shipped with zero test coverage in exactly the gap this test now closes.
What mocking alone could not catch — four real bugs, found only by testing against a live repo
Mocked tests describe the API the way the author believes it behaves. All four of these bugs passed a 100%-green mocked suite before being found:
- Silent field clobbering.
applyrebuilt the entirerequired_pull_request_reviews/required_status_checkspayload on any change, hardcoding three unmodeled GitHub fields (dismiss_stale_reviews,require_last_push_approval, status-checkstrict) toFalse— silently resetting them if a human had set them manually. Found during code review, confirmed with a live-repo reproduction before fixing. - Strict-mode phantom drift.
diff._SCHEMA_DEFAULTS["status_checks"]wasStatusChecksPolicy(required=[]), but the real API translators represent “no status checks configured” asNone— semantically identical, not==-equal. Everystrict: trueapply against an already-compliant, unconfigured branch reported permanent 1-field drift and issued an unnecessary API call, forever. Only surfaced by runningrepo-policy applywithstrict: trueagainst a real repository and noticing the tool claimed 1 change when nothing should have changed. No mocked test combined “strict mode” with “current state hasstatus_checks=None” — the exact combination that broke. __version__/PyPI version divergence.python-semantic-release’sversion_tomlconfig only updatespyproject.toml; the hardcoded string insrc/repo_policy/__init__.pysilently drifted across 4 releases. Found by literally runningpip install repo-policyin a clean venv and checkingrepo_policy.__version__againstpip show’s reported version.allow_fork_syncing’s wrong permissive default.diff._SCHEMA_DEFAULTS["allow_fork_syncing"]wasTrue, chosen to match the siblingrepo_securitytool’s own recommended baseline value — never independently verified against live GitHub. Only surfaced by runningrepo-policy applyagainst a real repository and independently checking the resulting branch protection viagh api: GitHub silently discardsallow_fork_syncing: trueon any branch wherelock_branchisfalse, resetting it tofalseregardless of what’s sent. BecauseTruewas also the valuefrom_api(None, ...)used to represent “nothing configured,”diff.resolve_desired()’s managed-scope current-state inheritance carried the broken pairing into any first-timeapplyagainst a previously-unprotected branch — even for apolicy.ymlthat never mentionsallow_fork_syncingat all. No mocked test could catch this: it requires a real GitHub API response to observe that a value sent in aPUTdoesn’t persist. Fixed by flipping the default toFalse(the value GitHub always honors regardless oflock_branch) and adding a model validator that rejects an explicitallow_fork_syncing: truedeclaration unlesslock_branch: trueis also declared — seeCHANGELOG.md’s “Correctallow_fork_syncing’s permissive default from true to false” and “Rejectallow_fork_syncing: truewithoutlock_branch: true” entries for the full fix history.
The pattern across all four: the bug was invisible to any test that only asserted repo-policy’s own internal consistency. Each one required checking repo-policy’s output against an independent, real source of truth — a live repo’s actual API state, or a real PyPI install.
Manual Verification Checklist (run before any release you don’t fully trust)
Steps 1–7 below are now automated in tests/e2e/test_fixture_repo.py (run via pytest -m e2e,
or the nightly .github/workflows/e2e.yml) — see CHANGELOG.md’s “Add full E2E lifecycle
coverage against the live fixture repo” and related entries for the build-out history. This
checklist remains as the human-readable reference and manual fallback for anyone without CI access
to the fixture repo.
Against a disposable repository:
validatea config, confirm exit 0 on valid / exit 2 on invalid.planagainst an unprotected branch — confirm it reports every field as a change.apply, then independently confirm viagh api repos/<owner>/<repo>/branches/<branch>/protectionthat the real state matches.planagain — confirm zero drift (real-world idempotency, not just the mocked test).- Repeat 2-4 with
enforcement: ruleseton a second branch, cross-checkinggh api repos/<owner>/<repo>/rulesets. - Remove a declared branch from the config with
strict: true,apply, confirm the orphaned ruleset (and only that one) is deleted. - Restore the repository to its original state.
The automated suite now also covers a scenario beyond these 7 steps: after apply, directly
disable the managed ruleset through the raw GitHub API (bypassing repo-policy), confirm audit/
plan report it as drift, repair it with another apply, and independently re-confirm the
branch’s live effective rules — see “Ineffective-ruleset repair” below.
E2E Test Isolation and Concurrency Safety
tests/e2e/ mutates one shared, persistent, real repository (shipsolid/repo-policy-e2e-fixture)
across every local run and every CI run. Two hazards follow directly from that: test-order
dependence within a single run, and two runs (a manual workflow_dispatch and the nightly
schedule, or two manual dispatches) racing each other’s mutations against the same repo. Both are
addressed structurally, not by convention:
- One collapsed lifecycle test, not several dependent ones. Earlier revisions of this suite
split “plan shows drift”, “apply succeeds”, “plan shows zero drift”, and “strict apply prunes”
into separate test functions that only produced a coherent story because pytest happened to run
them in file order against the same session-scoped
clean_fixture_repofixture — one docstring literally said “Must run after test_plan_reports_full_drift…”.test_full_policy_lifecycleintests/e2e/test_fixture_repo.pyreplaces all of them with one function whose only ordering is its own statement order: clean baseline →plan(full drift) →apply→ independent API verification → no-driftaudit/plan→ ineffective-ruleset repair (disable the managed ruleset directly via the raw API, confirmaudit/plansurface it as drift, repair with anotherapply, re-confirm the branch’s live effective rules) → strict prune (switch to a policy that drops the ruleset-enforced branch, confirm the orphan is deleted) → final cleanup. There is no longer a subset or reordering of this suite that can produce a different outcome. - Independent smoke tests stay independent.
test_validate_accepts_e2e_fixture_policies(pure config parsing, no API calls) andtest_live_token_and_repo_are_reachable(token/repo reachability only) never takeclean_fixture_repoand never assert anything about live policy state, so they remain correct regardless of what state the fixture repo happens to be in. - Cleanup is guaranteed even when setup itself fails partway.
clean_fixture_repo(tests/e2e/conftest.py) registers its teardown viarequest.addfinalizerbefore performing its first mutation, rather than relying on the tail half of ayield-based generator fixture. pytest only turns a generator fixture’s post-yieldcode into a registered finalizer once the code beforeyieldhas already returned successfully — so if a setup step after the first mutation had raised (e.g. the ruleset-branch existence check), the old design would never have reachedyieldand would silently skip scheduling any cleanup at all.request.addfinalizerhas no such gap: once registered, it always runs at session teardown, regardless of what happens afterward in this fixture’s own remaining setup or in any test built on top of it. - Concurrent workflow dispatch cannot overlap.
.github/workflows/e2e.ymldeclaresconcurrency: {group: repo-policy-e2e-fixture, cancel-in-progress: false}. A fixed group name (not templated on a ref or run id) means every trigger of this workflow — scheduled or manual — shares one queue;cancel-in-progress: falsequeues a second run behind the first instead of cancelling it, since a cancelled mid-mutation run would abandon the fixture repo in whatever partial state the cancellation caught it in, defeating the cleanup guarantee above. The net effect: GitHub Actions structurally never lets two runs of this workflow touch the fixture repo at the same time, regardless of how they were triggered.
Known Gaps
No CI-enforced coverage threshold— closed:ci.yml’stestjob runspytest --cov-fail-under=95on every Python version in the support matrix (3.10–3.14), andruff format --checkblocks merging on formatting drift.No automated end-to-end test against a real GitHub repository— closed: seetests/e2e/and.github/workflows/e2e.yml.- No test exercises GitHub’s classic-branch-protection-specific edge cases beyond what’s modeled
(e.g.
restrictionswith actual user/team push restrictions configured, ordismissal_restrictions/bypass_pull_request_allowances). 2026-09-22: attempted with a real second collaborator (shipsolid-release-bot) added toshipsolid/repo-policy-e2e-fixture— blocked by a GitHub API constraint, not a missing identity:422 "Only organization repositories can have users and team restrictions". The fixture repo is personal-account-owned (owner.type: "User"); GitHub rejects named user/team restrictions on any personal repo regardless of collaborator count. Live-verifying these three fields needs the fixture repo (or a second, dedicated one) to live under a GitHub organization — seeROADMAP.md’s “Next” table.