Time-Dependent Flaky Tests: Wall Clocks, Failures Under Load, Second Rounding and Fake Clocks
Why tests tied to the wall clock (fixed sleeps, short timeouts, duration assertions, single-instant "not yet" checks) fail precisely under load, how to measure that moment's pressure with PSI, how second rounding produces "too early" bugs, and how to fix a test without losing what it proves.
A test passes a hundred times on your laptop, fails in CI now and then, and passes again on retry. One of the most common causes of that kind of flakiness is time: the test carries an unwritten assumption about how long it takes to get from one line to the next, and the day the machine slows down, the assumption breaks. This post covers the patterns behind it, why they fail precisely under load, the "too early" bug that second rounding produces, and how to fix all of it without losing what the test proves.
Patterns that tie a test to the wall clock
Sleeping instead of synchronizing. Waiting for background work with time.Sleep(100 * time.Millisecond) claims the work finishes within 100 ms. A sleep is a lower bound; neither the sleep nor the work has an upper one.
Short timeouts and "must finish within N ms". Giving a 5 ms call a 100 ms budget looks generous locally. On a runner where dozens of processes share the same cores, the margin disappears. A unit test that asserts a duration measures the machine at that moment, not the code.
Checking "not yet" at a single instant. The sneakiest one:
func TestNotDispatchedEarly(t *testing.T) {
q := NewQueue()
defer q.Close()
q.Schedule("report", time.Now().Add(2*time.Second))
time.Sleep(time.Second)
if q.Dispatched("report") {
t.Fatal("job dispatched before it was due")
}
}
If the sleep actually takes two and a half seconds, the job was dispatched legitimately and the test blames the queue for a bug that does not exist.
Why they fail exactly under load
All of these assume the gap between two lines is small, and load stretches it. Under CPU starvation, parallel test packages and other jobs on the runner compete for cores, and a cgroup that exhausts its CPU quota is not scheduled at all for the rest of that period. Under memory pressure the kernel reclaims pages and allocating threads may stall. Another job writing to the same disk can hold up an fsync for seconds. Garbage collection pauses, and the CPU the collector burns, hurt more on a busy machine. The race detector is a multiplier of its own: according to the Go documentation, for a typical program it can increase memory usage by 5-10x and execution time by 2-20x.
One distinction matters: load also exposes real races in product code, so "it fails under load" does not mean "the test is wrong". Measure instead of guessing.
Measuring that minute's pressure with PSI
Since Linux 4.20 the kernel exports resource stalls as PSI (pressure stall information). In the cpu, memory and io files under /proc/pressure/, the some line is the share of time in which at least one task was stalled on that resource, and full is the share in which all non-idle tasks were stalled at once. Each line holds percentages over the last 10, 60 and 300 seconds plus a cumulative total in microseconds. These files describe the whole machine; a container's own pressure is in the cpu.pressure, memory.pressure and io.pressure files of its cgroup v2 group.
Read total at the start and end of a test and put the difference in the report when it fails (in Go, a t.Cleanup that checks t.Failed()). If failures line up with high pressure, the test is most likely time dependent. If it also fails on an idle machine, look elsewhere.
Second rounding and the "too early" bug
Math.floor(Date.now() / 1000), or a column type that drops fractional seconds, rounds time down. Round a "do not dispatch before" boundary down and the system can really dispatch up to 999 ms early: a job due at 12:00:00.700 is stored as 12:00:00, and the dispatcher treats it as due at 12:00:00.100.
The meaning of the boundary decides the direction. A "not before" boundary rounds up (Math.ceil), a "not after" boundary rounds down, so the error always lands on the safe side. Do not assume what the database does either: MySQL rounds to the nearest value by default when fractional seconds go into a column with fewer digits, so a fraction below one half still pulls the boundary down.
The same bug reaches tests from the other side. If the observed time has second resolution and the expected time has millisecond resolution, a job dispatched on time looks early on paper. Bring both sides to the same resolution before comparing. And measure durations with a monotonic clock: performance.now() rather than Date.now() in the browser; in Go, Round, Truncate and serialization strip the monotonic reading from time.Now().
Fix patterns
A fake clock. testing/synctest, stable since Go 1.25, hands time to the test without adding an interface: when every goroutine in the bubble is blocked, the fake clock jumps to the next event.
func TestDispatchBoundary(t *testing.T) {
synctest.Test(t, func(t *testing.T) {
q := NewQueue()
defer q.Close()
due := time.Now().Add(2 * time.Second)
q.Schedule("report", due)
time.Sleep(2*time.Second - time.Millisecond)
synctest.Wait()
if q.Dispatched("report") {
t.Fatal("job dispatched before it was due")
}
time.Sleep(time.Millisecond)
synctest.Wait()
if !q.Dispatched("report") {
t.Fatal("job not dispatched when due")
}
})
}
The test checks the boundary to the millisecond and never really waits. On older versions an injected Clock interface gives the same control.
Polling, and an assertion that does not depend on the clock. With the real clock, poll for the condition instead of sleeping, with a timeout a healthy run never reaches; that timeout is a guard rail, not part of the assertion. Instead of "one second later it must still not be dispatched", assert something that is always true: if the job has been dispatched, the clock read at that moment cannot be earlier than the due time.
func TestNeverEarly(t *testing.T) {
q := NewQueue()
defer q.Close()
due := time.Now().Add(200 * time.Millisecond)
q.Schedule("report", due)
deadline := time.Now().Add(30 * time.Second)
for {
given := q.Dispatched("report")
now := time.Now()
if given && now.Before(due) {
t.Fatalf("job dispatched %v early", due.Sub(now))
}
if given {
return
}
if now.After(deadline) {
t.Fatal("job not dispatched within 30s")
}
time.Sleep(10 * time.Millisecond)
}
}
Order matters: state first, clock second. The other way round, the clock can be read just before the due time, the job is dispatched in between, and the false failure is back. Under load this test never raises a false alarm; it only takes longer. A stronger version has the dispatcher record the time it read when handing out each job, and the test asserts !at.Before(due) for every record.
Counting instead of timing. The concern behind "under 50 ms" is usually an extra query or a network call in a loop. Assert the count and leave timing to benchmarks in a controlled environment.
Elements that render later. A menu that opens on click, animates in or loads its items asynchronously may be missing or re-rendered at the moment the test reaches for it. In Playwright, use locators, which are resolved again for every action, instead of a handle grabbed once, and scope the lookup to the open menu: page.getByRole('menu').getByRole('menuitem', { name: 'Refresh' }). Do not target by index, and do not skip actionability checks with force: true. Test time-dependent UI with a fake clock as well:
test('session warning does not appear early', async ({ page }) => {
await page.clock.install({ time: new Date('2026-01-05T09:00:00') });
await page.goto('/dashboard');
await page.clock.pauseAt(new Date('2026-01-05T09:10:00'));
const alert = page.getByRole('alert');
await page.clock.runFor('19:00');
await expect(alert).toBeHidden();
await page.clock.runFor('02:00');
await expect(alert).toBeVisible();
});
toBeHidden() is a single-instant "not yet" check too, but with the clock paused nothing can change in between.
Loosening a test versus fixing it
The first reaction to a flaky test is usually to loosen it: a bigger timeout, a longer sleep, a retry, a tolerance. If the timeout is only a guard rail, raising it costs nothing. When time is part of the assertion, loosening changes the assertion: a 500 ms tolerance on "the job is never dispatched early" lets through exactly the 400 ms early dispatch it was meant to catch. A longer sleep only lowers the odds of failure and slows the suite.
My rule: before the change, write down in one sentence what the test proves. After the change, check whether that sentence still holds. If it got weaker, the test was loosened, not fixed.
The risks of quarantine
A quarantined test protects nothing; if it was the only test covering a feature, the feature is now effectively untested. An entry with no owner and no deadline becomes permanent. More importantly, the cause is sometimes the product: a race that appears under load will appear for users under the same load, and silencing the test silences that signal. Automatic retries do the same less visibly; as long as one of three attempts passes, nobody notices.
If you need a quarantine, do not skip the test. Run it in non-blocking mode, record its results along with the PSI values at the moment it failed, and give every entry an owner and an end date. When the date comes, the test is either fixed or deliberately deleted. For most time-dependent tests the fix is one of the patterns above, and in my view it is cheaper than managing the quarantine list.