Skip to content

Borealis Testing Regressions

Track unit tests that protect prior regressions or currently fail during test formalization.

Status Labels

  • open: failure still reproduces.
  • fixed: test passes and protects the repaired behavior.
  • stale-test: product behavior changed and the test expectation needs updating.
  • environment-gap: test needs a dependency, runtime copy, OS feature, or service not present in the current runner.
  • flaky: test sometimes passes and sometimes fails without code changes.

Current Formalization Baseline

Baseline sampled on April 30, 2026 from branch feature/unit-test-formalization.

ID Area Test or Lane Status Notes Cleanup Rule
REG-TEST-001 Engine auth and Aegis Go auth_*_test.go, aegis_*_test.go, credentials_test.go, and TestWebUILeavesAuthCookieManagementToGoBackend fixed Go tests assert secure HttpOnly cookies, Aegis envelopes, credential storage, MFA, passkey, and password-reset behavior. Retired Python source scans were removed. Keep tests aligned with encrypted-at-rest auth storage. Do not reintroduce plaintext passkey, MFA, or browser-written auth cookies.
REG-TEST-002 Legacy Python Agent role loading Data/Engine/Unit_Tests/test_agent_role_manager.py stale-test Legacy Python RoleManager source and tests were removed after Go Agent parity. Do not reintroduce Python RoleManager coverage; add Go runtime role-registry tests under Data/Agent/internal when role wiring changes.
REG-TEST-003 Engine WebUI unit lane Engine_Unit_Tests.sh WebUI Vitest step fixed Runtime WebUI lane passes under bounded timeout. Stale Agent Health, route guard, and AppShell assertions were updated for current UI behavior. Keep WebUI tests under Unit_Tests/**/*.test.jsx and preserve bounded Vitest execution.
REG-TEST-004 Agent WireGuard role tests Data/Agent/internal/roles/wireguard_tunnel/wireguard_tunnel_test.go fixed Go WireGuard tests cover session normalization, Windows service naming, firewall port arrays, Linux wg-quick apply, missing wireguard-tools repair through Rocky/RHEL dnf, dependency retry cooldown, authenticated endpoint fallback, handshake-aware health/readiness, and signed token trust persistence. Keep WireGuard service/config/dependency/handshake/readiness tests aligned with real Windows and Linux tunnel behavior.
REG-TEST-005 Engine software inventory tests Data/Engine/Containers/api-backend/cmd/api-backend/software_test.go fixed Validation exposed tests depending on live icon override, uninstall override, and uninstall blocklist JSON. Keep override paths isolated with test-owned temporary files; load operator JSON only when a test explicitly covers hotload behavior.
REG-TEST-006 Engine RBAC and filters Data/Engine/Containers/api-backend/cmd/api-backend/device_filters_test.go fixed Full Engine validation exposed filter reload and usage-conflict checks running without current actor plus stale site-scope handling. Keep filter create/update/archive/delete visibility and usage contracts in Go API tests.
REG-TEST-007 Site-worker runner isolation Engine_Unit_Tests.sh site-worker Python lane fixed Single-process pytest run exceeded lane timeout and process-global worker state leaked between files. Runner executes current site-worker Python files in isolated pytest processes with per-file timeouts. Preserve all site-worker test execution while keeping file-level isolation and useful per-file logs. Go-owned behavior must stay in Go suites.
REG-TEST-008 Scheduled Ansible on K3s Data/Engine/Unit_Tests/test_ansible_runner.py::test_runner_bounds_concurrent_controller_processes and TestSchedulerInternalJSONHonorsRequestTimeoutBeyondSharedClientTimeout fixed Job 372 exposed 30-second HTTP cutoff inside 45-second WireGuard readiness window plus eight simultaneous Ansible controllers exhausting 512 MiB site-worker cgroup. Keep request-specific readiness deadlines authoritative and bound controller processes independently from scheduled work-item slots.
REG-TEST-009 Agent binary hot-publish and worker cutover TestEngineAgentBinaryRedeployKeepsOldWorkersUntilHealthCutover, test_stale_worker_exit_does_not_retire_replacement_registration, and TestBorealisOperatorLaunchSiteWorkerBuildsSafePod fixed LAB-TRAEFIK-01 redeploy reused stale Engine Agent cache. Go contract tests now protect readiness-first stable-Service cutover and pre-commit rollback ordering. Keep candidate readiness and Service health ahead of old-pod deletion. Preserve pre-commit selector rollback and incarnation-aware worker stop.
REG-TEST-010 Scheduled Ansible connection admission TestScheduledJobAggregationShowsConnectionProbeDeadline, TestScheduledConnectionProbeUsesSixtySecondWindow, Create_Job.connectionProbeStatus.test.jsx, TestVPNSessionPayloadIncludesEngineWireGuardFallback, TestEngineIPFallbackIsResolvedForEveryNetworkMode, and Agent WireGuard fallback/health tests environment-gap Jobs 380/381 and first two Job 382 runs skipped LAB-TRAEFIK-01 with no_handshake. Public-mode write_compose_env resolved Engine host fallback only inside Internal-Only CA branch, leaving deployed API fallback blank. Assignment now runs before profile-specific CA work. After API rollout, Agent selected 192.168.3.252:30000, reported fresh handshakes, and Job 382 run 47960 completed Success with Ansible pong, unreachable=0, and failed=0. Preserve persisted connection deadline, one-second UI derivation, non-wrapping status pills, lowercase seconds suffix, mutually exclusive live Job History status Filter Sliders, full-height target grid, 60-second SSH/WinRM window, active-run handling, public/local fallback emission, and handshake-proven VPN readiness. Restore deployed WebUI Unit_Tests, rerun WebUI lane, then mark fixed.
REG-TEST-011 Privileged WireGuard control boundary Data/Engine/Containers/api-backend/internal/wireguardcontrol/server_test.go fixed Python control proxy and its pytest were ported together. Go tests preserve exact listener, peer, route, and firewall command shapes and reject broad networks, Engine-address reuse, arbitrary commands, and symlink escapes. Keep privileged command allowlist narrow. Any new command shape needs direct allow and reject tests plus security-document update.
REG-TEST-012 Go route test evidence APIRouteEvidenceTests and Tests/policy/check_api_routes.py fixed PR review found route generator auto-assigned companion test files without proof author reviewed route coverage. Generator now preserves committed route-specific evidence and leaves new routes empty so policy fails until focused test or reviewed exemption is recorded. Never infer new-route coverage from companion filename alone.
REG-TEST-013 K3s hostPath boundary K3sHostPathPolicyTests fixed PR review found string-prefix comparison allowed sibling paths such as secrets-copy through secrets allowlist. Policy now accepts only exact directory or slash-delimited descendant. Preserve path-component boundary for every fixed hostPath prefix.
REG-TEST-014 Full portable container lane PortableRunnerContractTests fixed PR review found clean-worktree run-all.sh selected zero container images. Full runner now calls explicit run-containers.sh --all mode and builds every production manifest image. Keep worktree-aware default for narrow iteration; keep run-all.sh explicitly exhaustive.
REG-TEST-015 Empty affected-container selection AffectedServicesTests.test_cli_emits_no_empty_service_for_unknown_path fixed PR #445 exposed blank-line CLI output being parsed as an empty Docker service when workflow-only changes selected no images. Keep line-oriented service output empty when no image matches; preserve JSON and GitHub output contracts.
REG-TEST-016 Windows Server remote desktop readiness Agent VNC/RDP ready-fast-path, lock-expiry, coalescing, firewall-drift/postcondition, Windows single-host filter, failed-firewall readiness/cache invalidation, deferred-update reconciliation, and command-output tests; Engine VNC/RDP slow-device budget, non-auth RFB failure classification, and RDP transport-recovery tests; WebUI deadline test fixed Issue #438 showed Windows Server readiness spending 21-36 seconds behind serialized PowerShell/firewall work while Engine and browser stopped after 20-30 seconds. Agent updates also replaced binary/service without replaying full mutable host reconciliation. Initial August 20 validation then showed NetSecurity postcondition failures consistent with Agent requiring /32 text while Windows returned an equivalent single-host representation; VNC preserved stale readiness after that failure and could hand Guacamole credentials not yet applied to UltraVNC. After single-host normalization and cache invalidation deployed, LAB-DC-01 completed seven consecutive VNC and seven consecutive RDP operator connections successfully. VNC and RDP launches used cached listener fast paths; two logged RDP background audits completed ready in about 15 and 18 seconds without blocking launches or hitting 60-second budget. Preserve 60-second Agent readiness budget, 75-second total setup budget, install-equivalent Windows update repair with identity/trust preservation, exact-host firewall scope using Windows single-IP syntax with exact /32 read-back equivalence, cached matching listener fast path, five-minute heavy audit, context-aware foreground lock expiry, background coalescing, verified Borealis firewall mutation, failed-firewall cache invalidation, one bounded RDP transport recovery, PowerShell output diagnostics, and structured non-auth RFB banner stages without added VNCAuth attempts.
REG-TEST-017 Agent file-log retention TestApplyDefaultsMigratesLegacyLogRetention, TestRetentionDaysFromConfigMigratesLegacyDefault, TestRotateAndPruneKeepsOnlyConfiguredRetention, and TestRotateAndPruneDefaultKeepsSevenCalendarDays fixed Idle Agent role ensure diagnostics can produce tens of thousands of useful lines. Removing those diagnostics would weaken troubleshooting, while previous one-day retention discarded useful history too quickly. Agent logs continue daily rotation, retain seven calendar days by default, prune older rotated files on next write/start, and migrate previous generated one-day config value to seven days. Preserve ensure diagnostics and shared daily rotation. Keep seven-day default plus legacy one-day migration; retain explicit positive custom overrides.
REG-TEST-018 Install-equivalent Agent updates and progress dashboard Agent update persistence/health/ownership and Linux recovery tests, Engine agent_update_progress_test.go, heartbeat progress broadcast and maintenance timeout tests, WebUI Agent_Updates.test.jsx and workspace URL tests, plus operator-approved LAB-CA-01 same-build qualification fixed Issue #452 required operator updates to repair full install state even on same build, preserve identity/trust, expose durable Scheduler-backed progress, and keep no-update hourly checks non-disruptive. LAB-CA-01 operation 380adcc8-3632-43e4-b158-e709e66c85d1 retained current binary, skipped download and staging, reconciled managed components without restarting healthy native RDP, reconnected every Agent role, persisted Scheduler Job 396 history, and completed successfully in 40 seconds on August 20, 2026. PR review also found post-quiesce Linux staging failure could leave runtime stopped, terminal heartbeat replay caused repeated Scheduler writes and SSE invalidation, and timed-out runs left queued dispatch work claimable. Preserve pre-ack local operation persistence, exact Windows process ownership, post-quiesce Linux runtime restoration, terminal role-health gate, accepted-event-only Scheduler writes and SSE, queued-work cancellation on timeout, canonical deep links, and no same-build hourly restart/job.
REG-TEST-019 Quick Job Assembly provenance Quick_Job_Dialog.sourceDomain.test.jsx fixed Issue #326 found Quick Job scheduled components omitted Assembly domain metadata. Scheduled Job source badges could then classify Aurora Assemblies as User-Created before catalog enrichment or when catalog lookup was unavailable. Preserve selected Assembly domain and domainLabel in Quick Job component payloads. Missing catalog and saved provenance must display Unknown, never infer User-Created.
REG-TEST-020 Cluster controller lease and action Job restart safety TestClusterControllerLeaseGuardRenewsDuringLongStep, TestClusterControllerLeaseGuardCancelsStepOnOwnershipLoss, TestClusterControllerPostgresLeaseSerializesLongStep, TestNodeActionJobResumesMatchingAlreadyExistsCollision, and immutable Job identity tests fixed PR #462 three-node admission attempt outlived 20-second controller lease. Second Ready controller claimed same operation step and raced deterministic Job creation, turning matching 409 AlreadyExists into failed membership operation after workloads had already promoted. Keep independent lease renewal, cancellation on renewal/ownership failure, lease-fenced durable transitions, operation-attempt Job names, and fail-closed reuse of only exact matching action Jobs.
REG-TEST-021 Cluster Management truth and operation lifecycle TestClusterBannerIgnoresHistoricalFailures, TestClusterOperationHistoryMarksGlobalFailuresSuperseded, TestClusterDatabaseRuntimeStateTracksConfiguredAndReadyInstances, TestClusterDatabaseDegradationAllowsOnlyRecoveryMutations, TestClusterQuorumDegradationAllowsOnlyRecoveryMutations, TestCurrentReleaseAdmissionBatchSize, TestClusterCustomResourceStatesKeepDesiredAndRuntimeFieldsSeparate, and Cluster_Management.test.jsx fixed PR #462 operator screenshots showed later-successful membership still raising global failed banner, owner UUIDs, K3s version misclassified as role, empty release picker, unsafe retry on superseded failures, configured PostgreSQL count hiding one non-Ready CNPG instance, and emergency removal unable to complete or admit replacement because transient active size two was fenced as unsupported. Keep global banner limited to active states, immutable failed audit with transactional supersession fence, operator-facing node labels, explicit release empty state, live configured/Ready database counts, recoverable Degraded Database, recovery-only controls during degraded states, and exact 3 -> 2 -> 3 externally fenced replacement recovery without opening five-node membership.
REG-TEST-022 Rolling host node-manager activation TestNodeActionPodsActiveWaitsForRunningOrUnknownWork, TestWaitForActiveNodeManagerExecutableRequiresRunningInodeAndSocket, and test_cluster_redeploy_refreshes_host_node_manager_after_candidate_exists fixed PR #462 live rollout found clustered Engine update built target node-manager but only reconciled container workloads, leaving existing hosts on old root recovery code. Direct restart inside request would kill action and its Engine child. Preserve atomic binary replacement after candidate creation, shell-free transient activation outside service cgroup, two idle node-action observations, running-inode/socket verification, and durable target-SHA marker.
REG-TEST-023 Rolling edge/WireGuard role transfer TestEdgeRoleTransferAwayAcceptsDifferentHealthyEligibleOwner and TestRoleEligibilityDoesNotWaitForLegacyStandbyWireGuardReadiness fixed PR #462 three-node rollout showed kube-vip may elect any eligible survivor instead of controller's first replacement, and prior-release WireGuard readiness can fail after former owner withdraws interface. Exact replacement wait or pre-transfer standby readiness therefore deadlocked valid rolling update. Transfer-away succeeds only after edge lease leaves target and actual elected owner's WireGuard workload is ready. Keep exact-owner waits for HMR. Do not require fenced target or standby candidate readiness before edge election.
REG-TEST-024 Rolling drain endpoint withdrawal TestWaitNodeEndpointsWithdrawnIgnoresResidentInfrastructureEndpoints fixed PR #462 drains correctly removed API, WebUI, guacd, scheduler, Traefik, and site-worker traffic, but generic Borealis EndpointSlice waits first saw intentionally resident operator endpoints and later saw isolated API candidate endpoints required for Aegis key delivery. Both blocked progress after active traffic had withdrawn. Evaluate withdrawal only for active endpoints of Services controlled by application drain. Keep resident operator, database, and isolated candidate endpoints outside traffic-withdrawal gate while still rejecting any ready active target endpoint for drained traffic Services.
REG-TEST-025 Graceful VIP-owner host shutdown TestShutdownHandoffRunsOnlyWhileSystemIsStopping, TestShutdownHandoffParsersRequireReadyEnginePeerAndLeaseHolder, TestPerformShutdownHandoffWithdrawsEngineNodeAndWaitsForClusterVIP, TestPerformShutdownHandoffSkipsWithoutReadyEnginePeer, and K3s node-manager service policy fixed PR #462 edge-owner reboot stopped Traefik immediately while K3s KillMode=process left kube-vip containerd shim renewing dead VIP lease for 97 seconds. On whole-system shutdown only, temporarily withdraw local Engine-node eligibility and require fixed Cluster Virtual IP lease to move to alternate holder before K3s stops. Keep ordinary node-manager restart/update no-op, one-node shutdown unblocked, and wait bounded below systemd stop timeout.
REG-TEST-026 Transient workload drain marker K3s manifest transient-drain policy and TestBorealisOperatorLaunchSiteWorkerBuildsSafePod/transient_preStop_marker fixed PR #462 leader-owner partition restarted site-worker container inside existing Pod. Prior preStop left /tmp/borealis-draining in pod emptyDir, so replacement container inherited permanent readiness failure. Every readiness-withdrawal preStop must install cleanup trap before creating marker. Keep administrative maintenance/HMR drain in durable cluster and node-label state, and bump site-worker runtime contract when lifecycle template changes.
REG-TEST-027 Cluster mutation authentication TestClusterMutationsAcceptValidAdminSessionWithoutFreshStepUp and Cluster_Management.test.jsx fixed PR #462 operator qualification required repeated sign-outs and sign-ins because cluster mutations rejected otherwise-valid Admin sessions more than five minutes after access-token issue. Keep cluster mutations Admin-only and retain typed confirmations, operational gates, and actor audits. Do not reintroduce a cluster-specific short-lived step-up window inside a valid session.
REG-TEST-028 Shared Longhorn cluster storage TestSharedArtifactStorageExpandsAndVerifiesReplicaPerEngine, TestSharedArtifactStorageRejectsPVOwnershipMismatch, TestPlannedDisruptionRequiresSharedArtifactHAWhileRecoveryBypassesGate, Longhorn host-dependency policy, and test_longhorn_multipath_guard_* fixed PR #462 lost-HMR-target qualification found shared Agent artifacts had only one Longhorn replica on failed Engine. Recovery then exposed multipathd claiming Longhorn iSCSI device and preventing RWX share-manager remount with already mounted or mount point busy. Keep one healthy shared-artifact replica per three-node Engine, exact PVC/PV/Longhorn ownership checks, planned-disruption gate with recovery bypass, global one-replica StorageClass policy, genuine-multipath fail-closed behavior, and exact cleanup of Longhorn-owned maps only.
REG-TEST-029 HMR exit alternate VIP owner TestHMRExitRoleTransferAwayAcceptsHealthyEligibleOwner fixed PR #462 normal HMR exit moved PostgreSQL to first restored standby, but kube-vip validly elected other healthy standby for VIP/WireGuard ownership. Exact-owner waiting blocked exit even though production roles had safely left HMR target. Keep PostgreSQL placement exact. During HMR exit, accept any healthy non-target Cluster Virtual IP lease holder only after elected owner's WireGuard workload is Ready; HMR entry still requires selected target to own Cluster Virtual IP/WireGuard.
REG-TEST-030 Removed Engine node re-admission identity TestClusterMembershipAdmissionReactivatesRetainedNodeIdentity fixed PR #462 safe 3 -> 1 removal retained Engine 02/03 node rows for audit and recovery. Pair re-admission used new admission UUIDs with same unique node names, so completion looped on cluster_nodes_node_name_key after physical membership and workloads had already recovered. Reconcile admission completion by unique node name, preserve retained durable node ID and creation time, refresh mutable identity/runtime fields, clear removal drain reason, and mark new admission Admitted in same transaction.
REG-TEST-031 Cluster banner dismissal and event action safety AppShell.navigation.test.jsx and Cluster_Management.test.jsx fixed PR #462 operator validation found global cluster banner polling made banner effectively non-dismissible, while failed Cluster Events rows exposed inline retry capable of starting historical work outside deliberate recovery flow. Keep close action at far right and vertically centered, retain dismissal across polling for same banner identity, surface changed banner state again, and keep failed Cluster Events rows diagnostic-only without inline retry.
REG-TEST-032 HMR partition recovery and node-template convergence TestClusterControllerLeaseAcquireTimesOutStalePrimaryConnection, TestHMRLostTargetRecoveryFencesRejoinedTargetBeforeCommit, and test_reconcile_refreshes_node_clone_from_generic_template partial PR #462 all-role HMR partition qualification exposed three coupled gaps: standby controllers hung on stale CloudNativePG primary TCP connections, rejoined target could restore active labels before recovery committed its drain, and existing per-node Operator clone retained old site-worker image allowlist after generic template changed. Live reruns proved context deadline alone did not interrupt lib/pq socket read, then idle-pool purge could not reach a connection still blocked inside driver. Later scoped rebuild exposed old emergency kubectl set image field ownership and older clone selector shape blocking canonical convergence. Keep lease acquisition bounded and purge timed-out plus idle stale controller connections after the driver returns. Track active blackholed socket interruption in issue #466. Require rejoined-target drain/endpoint/edge-owner fencing before durable recovery commit. Derive per-node pod templates from canonical generic zero-replica Deployment while preserving existing immutable selectors, and let authoritative node-workload server-side apply reclaim conflicting generated fields.
REG-TEST-033 Rolling controller action-image transition TestNodeActionJobRejectsMismatchedExistingJob fixed PR #462 first published stable rolling update replaced controller image while source controller's exact PromoteCandidate Job still ran. Target controller acquired lease, requested same operation/attempt/step with its own action image, and failed otherwise-matching immutable Job before old Job exited successfully. Engine-update replay may accept only action-image difference when both references are immutable Borealis images and operation ID, attempt-derived Job name, full step, node, ServiceAccount boundary, command, and arguments remain exact. Every non-update image mismatch still fails closed.
REG-TEST-034 Cluster operation retry checkpoint TestClusterRetryResumesFailedCheckpointAfterPreflight and TestClusterOperationRetryPersistsFailedStepCheckpoint fixed PR #462 published rolling update failed with Engine03 drained. Existing retry reset operation to first node; it then drained Engine02 too and temporarily left only Engine01 serving application traffic. Persist failed step across retry and any retry-preflight failure. Re-run full preflight with fresh attempt identity, then resume exact valid checkpoint without replaying already-completed nodes. Invalid checkpoint fails closed.
REG-TEST-035 Rolling node-action image availability TestClusterControllerClaimPinsOperationActionImage and TestClusterOperationActionImageFallsBackBeforeClaimPin fixed PR #462 controller promotion changed BOREALIS_CLUSTER_ACTION_IMAGE to target API image before later Engine nodes staged that image. Subsequent target-pinned action Job could remain ErrImagePull, as Engine01 did during recovery attempt 2. First controller claim must persist current immutable action image in operation payload. Every controller holder uses pinned source image for whole attempt and later retries; target image becomes controller default only for new operations after rollout.
REG-TEST-036 Cluster node-action Job API restart tolerance TestTransientKubernetesAPIErrorAcceptsEOF and TestWaitJobToleratesTemporaryKubernetesAPIOutage fixed PR #462 .3 rolling update completed Engine03 drain Job, but controller failed operation when Kube API closed Job GET response with bare EOF during service restart. Existing classifier covered unexpected EOF, HTTP 429/5xx, connection refusal, and connection reset, but not wrapped io.EOF. Preserve wrapped io.EOF classification as transient while keeping authorization and TLS trust failures terminal. Keep operation context as polling bound.
REG-TEST-037 Rolling generic template image convergence test_promotion_refreshes_zero_replica_generic_images_after_active_health, test_generic_image_refresh_rejects_serving_template, test_generic_image_refresh_requires_resource_version, test_generic_image_refresh_rejects_stale_patch_response, and test_generic_image_refresh_rejects_unexpected_patch_container fixed PR #462 .29.1 rolling qualification promoted every node-scoped workload to target image while generic zero-replica API and cluster-controller templates retained prior stable image references. Active traffic remained correct, but canonical template state did not converge with promoted release. Refresh matching generic container images and revision only after active rollout health and before candidate deletion. Require generic template to exist at exactly zero replicas, require exact immutable container-name/image map, guard patch with fetched Kubernetes resourceVersion, validate returned state, and fail operation on invalid state, concurrent mutation, or patch failure.
REG-TEST-038 Maintenance and isolation VIP placement TestReconcileVIPPlacementLabelsNodesBeforeApplyingClusterSelector, TestVIPRoleTransferWaitsForClusterLeaseAndWireGuardReadiness, TestMaintenanceAndHMRReconcileVIPPlacementBeforeRoleTransfer, and kube-vip manifest policy fixed Issue #473 showed old Control VIP DaemonSet remained eligible on application-drained nodes because only Edge VIP had role-eligibility selector. Maintenance and Cluster-Wide Node Isolation could therefore leave K3s VIP ownership on drained standby. Use one Cluster Virtual IP DaemonSet requiring both existing eligibility labels. Initialize both labels from application state before selector migration, persist joined-node eligibility, require one lease to leave maintenance target, and require same lease on exact HMR target. Keep embedded-etcd membership independent from application maintenance.
REG-TEST-039 Fresh Engine bootstrap ordering and ownership test_debian_longhorn_dependency_install_waits_for_apt_lock, test_longhorn_probe_guard_waits_for_csi_daemonset_creation, test_sudo_checkout_owner_requires_matching_nonroot_account, test_repo_sync_reowns_source_but_prunes_runtime_state, and Longhorn host-dependency policy fixed Fresh Engine 01 redeploy first encountered unattended-upgrades holding dpkg lock, then failed because Longhorn driver-deployer reported Available before controller-created longhorn-csi-plugin DaemonSet existed. Root-run repository sync also left checkout unusable by invoking operator for normal Git work. Fresh 2026.09.1 qualification later found deployment-time Git inspection could rewrite .git/index as root after initial ownership repair. Keep bounded apt lock wait, CSI creation-before-patch gate, second dynamic rollout inventory, validated sudo-origin checkout ownership before and after successful command dispatch, and strict Engine/Engine.old/Agent runtime ownership pruning.
REG-TEST-040 Fresh WireGuard owner bootstrap TestBootstrapRuntimeCreatesIdentityAndCompleteListener, TestReconcileRuntimeLeavesCleanStandbyUnconfigured, TestWithdrawalRequestSuppressesOwnershipReactivation, test_wireguard_rollout_failure_collects_pod_diagnostics, test_wireguard_runtime_directories_allow_control_group_writes, test_wireguard_retry_resets_stale_progress_deadline, and WireGuard control bootstrap policy fixed Fresh Engine 01 deploy started K3s wireguard-tunnel before API. Control process created Unix socket but API alone generated server keys and activated borealis-wg, while WireGuard readiness required live interface/listener/route. First bootstrap fix then exposed root control Pod lacked DAC_OVERRIDE and could not create files inside 0750 borealis-engine:borealis-engine directories despite Borealis primary group. After permission correction, old Pod recovered on next CrashLoop retry but Deployment retained ProgressDeadlineExceeded, making immediate redeploy fail before observing recovery. Control process owns base listener bootstrap before readiness, API reuses generated identity and owns peer state afterward, standby remains withdrawn, preStop cannot race ownership reactivation, and failed rollout retains Pod plus host diagnostics. Keep only WireGuard config/secret directories 0770 for shared service-group creation while contained files remain 0640; do not add broad filesystem capability. Hash permission contract and reset unchanged failed rollout before waiting.
REG-TEST-041 Public certificate readiness behind outer proxy test_public_traefik_uses_tls_alpn_challenge and test_public_tls_readiness_requires_trusted_hostname_certificate fixed Fresh public Engine deployment marked WebUI accessible after Traefik ping while HTTP-01 challenge was intercepted by outer proxy and returned 502. Traefik stayed available with default self-signed certificate, breaking Agent onboarding trust. Use TLS-ALPN-01 through required public TCP 443, require outer proxy TLS/ALPN pass-through, and wait for locally served hostname certificate to pass normal CA validation before standalone deploy reports WebUI accessible. Append Traefik log diagnostics on timeout.
REG-TEST-042 Minimal first-node cluster identity TestClusterEnableDerivesAMD64NodeIdentityAndSynchronizesVIPStorage, TestClusterJoinRejectsNonAMD64ArchitectureBeforeAuthentication, TestClusterVirtualIPRequiresPrivateIPv4, Cluster_Management.test.jsx, and kube-vip manifest policy fixed Fresh cluster enable dialog required operator to enter two VIPs plus values Borealis already knows: current node management IPv4, node name, and architecture. Architecture selector also advertised unsupported ARM enrollment. Accept one private cluster_vip, derive first-node identity from Kubernetes Pod metadata and runtime architecture, reject non-AMD64 joins, synchronize legacy VIP columns during storage transition, and use one kube-vip lease/owner for K3s API, ingress, and WireGuard.
REG-TEST-043 Unreleased first-cluster baseline TestClusterEnableRejectsRetiredTypedConfirmationField, TestValidClusterBaselineReleaseRequiresDevelopmentNameToMatchSHA, TestClusterDevelopmentBaselineCatalogStopsAfterFirstPageAndSelectsApprovedChannels, TestClusterDevelopmentBaselineAcceptsApprovedQualificationPrerelease, TestHMRPinnedRestoreUsesSavedImmutableRelease, TestValidPinnedReleaseAcceptsStableQualificationAndMatchingDevelopmentIdentity, TestVerifyReleaseRefFetchesCommitBackedDevelopmentIdentity, test_engine_release_version_uses_clean_commit_backed_development_identity, and Cluster_Management.test.jsx fixed Fresh feature-branch deployment could not enable first cluster because API required published dotted-numeric GitHub release and redundant ENABLE CLUSTER text. No cluster-capable release could exist before cluster behavior completed qualification. Development baseline then made release catalog search every page for matching current tag that cannot exist. Enable request accepts only cluster_vip. Prefer exact stable or qualification tag, otherwise derive dev-<first-12-commit-characters> only from clean checkout and bind it to full SHA through API, controller, CRD, node manager, membership, and HMR. For development baseline, inspect one newest release page and allow only API-approved immutable stable or qualification descendants. Keep probe conformance mandatory.
REG-TEST-044 First-cluster WireGuard identity projection TestEnsureServerKeysAcceptsProjectedSecretSymlinks, TestReadWireGuardKeyRejectsSymlinkOutsideKeyDirectory, TestWireGuardRuntimeCreatesGroupWritableKeys, TestWireGuardRuntimeCreatesGroupWritableConfig, TestWireGuardRuntimeRepairsSharedConfigModeBeforeExistingListener, test_wireguard_runtime_directories_allow_control_group_writes, and WireGuard control bootstrap policy fixed Operation 13468bb5-83d1-4783-be75-8d8e11b6f979 converted first node, then per-node WireGuard workload rejected normal Kubernetes projected Secret symlinks as non-regular keys. Failed handoff restored generic workload, where API tunnel manager reset shared config directory from 0770 to 0750; capability-minimized root control process had Borealis group but could not create atomic config. Resolve projected key target and accept it only when regular file remains inside mounted key directory. Reject external symlink targets. Keep API manager, control bootstrap, and Engine ownership contract aligned on 0770 shared config/key directories with 0640 files, including repair before reusing existing listener; do not add DAC_OVERRIDE.
REG-TEST-045 Development release ancestry and picker bounds TestClusterDevelopmentBaselineCatalogStopsAfterFirstPageAndSelectsStableRelease, TestClusterDevelopmentBaselineRejectsStableReleaseOutsideAncestry, and Cluster_Management.test.jsx fixed PR #479 review found development baseline skipped source-ancestry compatibility, allowing unrelated stable tag into update submission before node manager rejected it. Operator screenshot also showed full-width release selector expanding across almost entire viewport and consuming most available height. Follow-up review found UI described security calendar versions as YYYY.MM.DD.N and exposed older incompatible releases. Pass full development baseline SHA through catalog cache and hydration, require GitHub compare status ahead or identical, fail closed when ancestry cannot be verified, and keep release menu bounded. Describe authoritative YYYY.MM.REVISION[.HOTFIX] policy, hide numeric downgrades, and show development-baseline targets only after API ancestry and compatibility approval.
REG-TEST-046 Probe conformance API cache invalidation K3s API manifest render and policy contract fixed PR #479 retained-host qualification passed all ten K3s probe trials, but same-image production redeploy left API Pod template on earlier failed status because bridge config hash omitted probe result and K3s version. Cluster enablement remained incorrectly gated. Capture probe status and K3s version once per reconciliation, include both in API bridge config hash, render those exact values into Pod environment, and bump bridge schema whenever cache inputs change.
REG-TEST-047 Immutable cluster qualification channel TestClusterDevelopmentBaselineCatalogStopsAfterFirstPageAndSelectsApprovedChannels, TestClusterDevelopmentBaselineAcceptsApprovedQualificationPrerelease, TestClusterQualificationUpdateRequiresWholeClusterAndExplicitConfirmation, TestCompareClusterReleasesOrdersQualificationBeforeStable, TestQualificationUpdateDefersSchemaFinalization, TestValidPinnedReleaseAcceptsStableQualificationAndMatchingDevelopmentIdentity, test_engine_release_version_uses_clean_commit_backed_development_identity, and Cluster_Management.test.jsx fixed Stable-only selector forced cluster test changes either through mutable development isolation or full supported release publication. GitHub prerelease had no channel contract, persistent support warning, or safe schema-promotion boundary. Accept only YYYY.MM.REVISION[.HOTFIX]-rc.N with matching GitHub prerelease flag and manifest permission. Require full-cluster DEPLOY QUALIFICATION, persist channel plus last stable identity, defer contract schema phase, reject downgrades/unrelated ancestry, and promote forward through normal stable release.
REG-TEST-048 Exact stable Engine release bootstrap test_release_packaging_avoids_admin_only_settings_api, remaining Tests/Unit_Tests/test_engine_release_bootstrap.py, and publish-engine-release-assets.yml workflow policy fixed Production curl instructions executed mutable main/Engine.sh; stable channel then selected latest numeric tag and fell back to main when tag lookup failed. Release assets lacked machine-verifiable Engine source identity. First live packaging run also called admin-only immutable-release settings endpoint with normal workflow token and failed 403. Fresh Ubuntu 24.04 AMD64 qualification installed immutable 2026.09.1 at commit d74cb92dcf8f3f037c33289fe661779af7ce7340; one K3s node, Borealis workloads, Longhorn PVCs, API healthcheck, trusted local HTTPS, and HTTP redirect all passed. Keep file-based bootstrap, immutable published non-prerelease release requirement, GitHub and manifest digest checks, exact full tag SHA, draft-first asset publication, and stable sync with no latest lookup or mutable fallback. Maintainer verifies repository setting before dispatch because workflow token lacks Administration permission; bootstrap enforces immutability after publication. Mutable refs remain explicit unstable development path.
REG-TEST-049 HMR production WebUI restoration test_webui_candidate_drops_hmr_development_runtime_state fixed Single-node HMR exit staged pinned production WebUI image from generic zero-replica template still configured for development. Candidate inherited BOREALIS_WEBUI_MODE=dev plus source mounts, then failed with vite: not found because production image contains served assets instead of development toolchain. Every WebUI release candidate must force production mode and remove all HMR-only source, configuration, test, and Vite cache volumes and mounts before apply. Keep non-candidate HMR workload unchanged.
REG-TEST-052 Aegis certificate renewal and independent trust TestAegisClusterTLSRenewalPartialProjectionAndExpiry, TestAegisClusterTLSCAOverlapAndRetirement, TestAegisClusterTLSRejectsMixedGenerationAndConcurrentReload, TestAegisClusterPostgresRenewedHolderUnlocksJoiningReplica, and Aegis workload trust tests fixed Long-lived API listeners kept startup certificates after cert-manager renewal, eventually preventing verified memory-key propagation. Leaf Secret CA replacement also lacked explicit overlap and retirement policy. Fresh TLS handshakes adopt validated projection snapshots, retain only still-valid credentials under latest accepted trust, and reject expired or retired identities. Preserve independent public trust across redeploys; keep one unlocked holder during staged CA rotation. Real PostgreSQL/mTLS tests prove joining-replica verification and locked cold restart; exact-release three-node qualification remains required under #493.

| REG-TEST-050 | PostgreSQL coverage and required-result enforcement | Tests/Unit_Tests/test_postgres_inventory.py and PostgreSQL inventory/result audit | fixed | Clustering review #493 found four existing maintenance, completion, retry-checkpoint and action-image integration tests omitted from runner, plus ordinary cluster Go changes missing database CI selection. | Discover all database-gated Go tests, compare maintained inventory, select every required test, reject skips and incomplete results, and select database lane for API/store/controller/scheduler source changes. Preserve per-test JSON results in CI. |

| REG-TEST-051 | Identity-bound planned member removal | cluster_removal_fence_test.go, member_removal_test.go, and TestManagedEtcdRemovalWaitsForK3sConfirmation | fixed | H01 review under #493 found NotReady/missing Node treated as fence proof, including retry storage shortcuts. Network partitions can produce those observations before any host fence. | Persist target/operation/Node UID/etcd identity before fixed host action; acknowledge only matching success under live controller lease. Require proof before membership mutation; reject partitions, stale operation replay, hostname reuse and changed member identity. Recover interrupted acknowledgement only from exact completed Job and prior intent. PostgreSQL test is required in H07 inventory. |

| REG-TEST-053 | Immutable release and authoritative K3s identity | cluster_release_identity_test.go, TestClusterReleaseIdentityPostgresQueueAndRetry, TestVerifyReleaseRefFetchesCommitBackedDevelopmentIdentity, and test_engine_redeploy_preserves_newer_installed_k3s | fixed | H06 review found cluster update accepted mutable GitHub releases, fetched manifests through movable tags, and reused picker compatibility calculated from stale API K3s environment after upgrade. | Refresh immutable publication for queueing, resolve one SHA and fetch its manifest, retain verified identity across retry, reject source changes under database lock, and observe K3s before mutation. Preserve installed K3s during Engine redeploy and qualify the exact manifest baseline in Q01. |

| REG-TEST-054 | Clustered HMR entry gate and retained restoration | cluster_hmr_gate_test.go, TestClusterHMRStartRejectsEntryBeforeStore, test_dev_dispatch_blocks_before_runtime_preparation, CLI membership/recovery tests, Cluster_Management.test.jsx, and retained #492 production candidate tests | fixed | U01 review requires disabling clustered development while existing isolation remains recoverable. CLI membership lookup failures and missing Kubernetes CRs during conversion previously appeared standalone; alternate dev service actions could rewrite shared WebUI mode. Legacy queued entry could bypass current API controls; cancellation erased its recovery target. Race CI also exposed unsynchronized request counting in the adapted storage fixture. | Reject entry before API/store/controller runtime mutation and all dev dispatch; require independent PostgreSQL proof when Kubernetes cluster resources are absent, and distinguish confirmed standalone from unavailable membership. Preserve pinned exit/retry and cancellation recovery state, recorded target and production mount cleanup. Wait for timed-out fixture handlers before reading request counts. Qualify exact-release restoration before deploying disablement. |

| REG-TEST-055 | Per-node Engine version visibility and honest freshness | Cluster_Management.test.jsx and TestClusterSnapshotPreservesNodeVersionRecordsAndPendingTarget | fixed | #476 and U02 review found Nodes grid omitted individual releases; recorded metadata could be confused with live runtime identity, and quiet failed polling retained apparently current data. | Show recorded release and full SHA, unknown identity, report age and stale snapshot. Keep pending target separate and scoped to active Engine-update targets. Preserve mixed records and refresh full-SHA tooltips even when release text stays unchanged. Recover after failed/hanging polls, preserve loader/poll snapshot receipt before auxiliary waits, accept sustained slow responses and independent catalog/history completions during snapshot failure, and reject superseded completions without claiming runtime version verification. Q01 verifies all three live nodes. |

| REG-TEST-057 | Agent enrollment binary validation | TestAgentEnrollmentRequestHandlerCreatesPendingResponse and TestEnrollmentBase64InputPreservesEncodingAndRejectsMalformedValues | fixed | H03 CI rejected a valid random Ed25519 SPKI as executable markup when base64 matched an event-attribute suffix. Deterministic key/nonce fixture reproduces failure. | Preserve bounded base64 protocol fields and existing decoding semantics; reject malformed, oversized and control-bearing values while retaining ordinary text markup checks. Keep deterministic fixture; do not retry flaky enrollment tests until green. |

| REG-TEST-058 | Fresh cluster node preparation | test_join_preparation_provisions_identity_and_build_dependencies and test_cluster_runtime_hydration_precedes_fresh_preparation | fixed | Admission on fresh hosts exposed absent runtime identity/build dependencies, Python bootstrap ordering and runtime preparation before authoritative configuration hydration (#511/#513). | Prepare account and OS dependencies before sandboxed deployment, then resolve FQDN/shared credentials and the CloudNativePG endpoint from cluster configuration in both environment files. Retain node-local paths, private file modes, literal secret syntax and failure cleanup; restore previous compose/runtime bytes and metadata after early or partial-render failure; verify complete live admission separately. |

| REG-TEST-059 | Immutable legacy admission configuration recovery | test_legacy_admission_recovery.py and retained test_engine_cluster_recovery.py preparation tests | live qualification open | Failed legacy admission pins pre-F03 Engine source while quorum state blocks normal update (#517). | Verify immutable repair source, original operation/cohort/node identity, drained roles, healthy K3s/CNPG and idle action state before target-local configuration repair. Reuse F03 transaction without runtime staging. Preserve prior bytes/metadata, stop writing children on timeout, restore target node-manager, reject symlinks and changed observations, retain private journal and resume through original controller gates. |

Detailed Codex Breakdown
  • Unit Testing
  • Engine Runtime
  • Agent Runtime
  • Security Whitepaper

  • Add a row when introducing a regression test for a bug found in production, operator validation, or a PR review.

  • If a formalization run exposes a failure that predates the cleanup, record it here before changing or quarantining the test.
  • Prefer fixing stale expectations over marking tests skipped.
  • If a skip or xfail is necessary, include the regression ID in the marker reason.