要点: 分散ReliabilityはHappy-path E2E追加ではなく、制御BoundaryでAssumptionを壊して検証します。
課題
環境Topology差、Flaky E2E、MockのProtocol漏れ、Retry・Clock・Partial Deploy時だけのFailureがあります。
単純な対策だけでは不十分な理由
巨大UI SuiteでなくContractとReal Infrastructure Component Testを使います。
ChaosはSafety Control付きExperimentです。
処理フロー
Compatibility→Component Test
Real Dependency→Failure Injection
Partial Fault→Production Verify
Canary・SLO
アーキテクチャ上の判断
Network下のInvariant Test
State MachineとIdempotencyをProperty Testします。
Ownership Boundary Contract
Error・Compatibilityも検証します。
Progressive Production Verify
Shadow、Canary、Flag、SLO Gateを使います。
さらに深く考える
Test DataもArchitecture
代表Dataを安全に用意します。
Game DayはPeopleもTest
Role、Runbook、Communicationを訓練します。
実装手順
- State TransitionとInvariantをDeterministicにTestします。
- 実Serialization FixtureでContractを検証します。
- Failureを注入しProduction SLOで判断します。
技術例: Controlled Chaos Experiment
hypothesis: checkout SLO remains healthy when one payment replica fails
scope: 5% internal canary traffic
inject: terminate one replica for 10 minutes
abort: error budget burn > threshold or queue age > limit
observe: retries, bulkhead, p99, duplicate charges
record: result, gaps, owner actionsFault、Scope、AbortをReviewed Automationで再現可能にします。
想定すべき障害パターン
- Mockが実Providerと違います。
- RetryがFlakyを隠します。
- 別Incident中にChaosします。
監視すべきシグナル
| シグナル | 確認する理由 |
|---|---|
| Flaky・Retry-to-pass | Suite信頼性です。 |
| Contract Failure | Deploy前Driftを検出します。 |
| Canary SLO・Rollback時間 | Production Verify効果です。 |
設計の検証方法
- 実SchemaでFixtureを検証します。
- Duplicate、Delay、Timeout等を注入します。
- Known-bad CanaryのPause・Rollbackを測ります。
安全なRollout計画
RiskをTest LayerへMapしFlakyを先に安定化します。内部Experiment、Small Canaryの順で拡大します。
本番前チェックリスト
- Critical DependencyにFailure Testがあります。
- ChaosにHypothesis、Scope、Abort、Ownerがあります。
- CanaryとIncidentが同じMetricを使います。
まとめ
Test環境はProductionを証明できないため、Changeを小さく観測・Rollback可能にします。
