From ce
Diagnoses and fixes tests that pass in isolation but fail when run concurrently. Covers shared state isolation, resource conflicts, and timing-based flakiness.
How this skill is triggered — by the user, by Claude, or both
Slash command
/ce:fixing-flaky-testsThe summary Claude sees in its skill listing — used to decide when to auto-load this skill
If the current repo has its own rules/skills covering this topic (check .claude/rules/ and repo CLAUDE.md), those take precedence — apply this skill only where they're silent.
If the current repo has its own rules/skills covering this topic (check .claude/rules/ and repo CLAUDE.md), those take precedence — apply this skill only where they're silent.
Target symptom: Tests pass when run alone, fail when run with other tests.
Test passes alone, fails with others?
│
├─ Same error every time → Shared state
│ └─ Database, globals, files, singletons
│
├─ Random/timing failures → Race condition
│ └─ See async waiting patterns in `writing-tests` skill
│
└─ Resource errors (port, file lock) → Resource conflict
└─ Need unique resources per test/worker
Quick diagnosis:
Tests pollute state that other tests depend on. Fix by isolating state per test.
| State Type | Isolation Pattern |
|---|---|
| Database | Transaction rollback, savepoints, worker-specific DBs |
| Global variables | Reset in beforeEach/afterEach |
| Singletons | Provide fresh instance per test |
| Module state | jest.resetModules() or equivalent |
| Files | Unique paths per test, temp directories |
| Environment vars | Save/restore in setup/teardown |
Database isolation (most common):
# Python: Savepoint rollback - each test gets rolled back
@pytest.fixture
async def db_session(db_engine):
async with db_engine.connect() as conn:
await conn.begin()
await conn.begin_nested() # Savepoint
# ... yield session ...
await conn.rollback() # All changes vanish
// Jest: Reset mocks between tests
beforeEach(() => {
jest.clearAllMocks()
jest.resetModules() // Clear module cache before test
})
afterEach(() => {
jest.restoreAllMocks() // Restore spied functions
})
See language-specific references for complete patterns.
Tests don't wait for async operations to complete.
See the writing-tests skill for async waiting patterns:
findBy, Playwright auto-wait)Quick summary: Wait for conditions, not time:
// Bad
await sleep(500)
// Good
await waitFor(() => expect(result).toBe('done'))
Multiple tests or workers compete for same resource.
Worker-specific resources:
# Python pytest-xdist: unique DB per worker
@pytest.fixture(scope="session")
def database_url(worker_id):
if worker_id == "master":
return "postgresql://localhost/test"
return f"postgresql://localhost/test_{worker_id}"
// Jest/Node: dynamic port allocation
const server = app.listen(0) // OS assigns available port
const port = server.address().port
File conflicts:
import tempfile
@pytest.fixture
def temp_dir():
with tempfile.TemporaryDirectory() as d:
yield d
| Stack | Reference |
|---|---|
| Python (pytest, SQLAlchemy) | references/python.md |
| Jest / Testing Library | references/jest.md |
| Playwright E2E | references/playwright.md |
After fixing, verify the fix worked:
# Run the specific test many times
pytest tests/test_flaky.py -x --count=20
# Run with parallelism
pytest -n auto
# Jest equivalent
jest --runInBand # First verify serial works
jest # Then verify parallel works
npx claudepluginhub rileyhilliard/claude-essentials --plugin ceDiagnoses non-deterministic test failures and eliminates root causes (timing, shared state, concurrency, external dependency, randomness) instead of retrying or skipping.
Diagnoses flaky tests by running them in a loop and identifying common causes (timing/ordering, async races, resource leaks). Use when CI fails intermittently.
Diagnoses and eliminates flaky or nondeterministic tests by classifying failure types (ordering, timing, resource, environment, external, concurrency) and isolating root causes with reproducible fixes.