What Is a Race Condition and How to Avoid It (2026)

What is a race condition and how to avoid it? A race condition is a defect in concurrent software where the result depends on the unpredictable timing or ordering of operations between two or more threads or processes touching shared data. You avoid it by making shared operations atomic, protecting them with locks or transactions, and giving every piece of state one clear owner.

If you write multithreaded code, run background jobs, build web APIs, or automate anything with shell scripts and cron, you will eventually meet one. Most of the time it shows up as a counter that came out short, a file that got clobbered halfway through, or a webhook that processed the same payment twice.

The awkward part is that it usually passes every test you have. Below is what the pattern actually is, what it looks like in real code, and which fixes hold up under load.

Table of Contents

What Is a Race Condition?

What Is a Race Condition?

A race condition is a defect in a program where the output depends on the uncontrollable timing or sequence of events between multiple threads or processes. When two of them touch the same shared resource without coordination, the result is non-deterministic: run the same program twice and you can get two different answers.

In simple terms, two people are reaching for the last item on a shelf at the same time. Both check that one item is still there, both decide to take it, and now there are two winners and one item. Nothing crashed. The program behaved exactly as written, which is what makes this bug so awkward to argue about.

The window between the check and the take is called the race window. Shrink it and the bug becomes rarer, but rare is not fixed. Attackers do not care how narrow it is, because they can fire hundreds of requests at it.

You may also see the older term race-around condition used for the narrowest case: an unprotected shared variable that every thread increments or decrements, fixed with atomic instructions. It is a special case of the wider problem, not a separate phenomenon.

Race conditions are not the only concurrency bug. A deadlock stops progress entirely while a race condition produces a wrong answer that keeps moving. Starvation is in between: one thread never runs because others keep taking the resource, and the system is alive but that work never finishes.

How Do Race Conditions Happen?

Every race needs three things at once: a shared resource, at least two actors, and no ordering guarantee between them. Remove any one and the bug disappears.

The classic case is check-then-act. Your code verifies a precondition and then acts on it:

if cache.count < 10:
    cache.count = cache.count + 1
    issue_ticket()

Between the comparison and the assignment, another thread can run the same block. Both threads see count at 8, both set it to 9, and only one ticket is issued while the counter says 9. That silent mismatch is a lost update.

Incrementing looks like one instruction, which is why it slips through code review. Underneath it is three separate steps:

read   count        // load the current value into a register
modify register    // add 1 to the register
write  count        // store the register back to memory

Here is the interleaving that loses the update. Two threads each run the loop body once:

T1: read count = 100
T2: read count = 100
T1: modify -> 101
T2: modify -> 101
T1: write count = 101
T2: write count = 101

Final value: 101, not 102. One increment vanished.

Multiply that by threads and iterations and the shortfall gets dramatic. A well-known demo that runs ten threads for ten thousand increments each frequently prints a number like 61505 instead of 100000, purely because of how the scheduler happened to interleave.

Shared memory is the obvious place to look, but it is not the only one. Races also show up in files two scripts rewrite, database rows two requests update, device or driver state, and external systems such as payment gateways and package registries.

What causes a race condition?

  • Shared state that more than one thread or process can modify without coordination.
  • Non-atomic operations built from several steps, such as a read-modify-write sequence or a check-then-act block.
  • Lack of synchronization, meaning nothing enforces ordering between the competing operations.

The most exploitable variant is TOCTOU, short for time-of-check to time-of-use. The program validates something and then uses it later, and the gap between those two actions is an attack window. OWASP tracks the general form as CWE-362 and the check-then-use variant as CWE-367.

A file permission check is the textbook case: you confirm a path is a regular writable file, then open it. In between, an attacker with local access swaps it for a symlink pointing somewhere sensitive. The fix is to not perform two steps at all. Open the file with the protection flags you need and inspect the handle you get back, rather than checking first and opening second.

What Does a Race Condition Look Like in Code?

Once you know the shape, you recognise it fast. Here is the unsafe version most people start with, in Python:

counter = 0

def increment():
    global counter
    counter += 1

threads = [threading.Thread(target=increment) for _ in range(100)]
for t in threads:
    t.start()

The single line counter += 1 is three operations. Fix it with a lock around the smallest possible region:

lock = threading.Lock()

def increment():
    with lock:
        global counter
        counter += 1

The shorter fix is to stop splitting the operation at all. In CPython, itertools.count() hands out values from a single indivisible step, so next() never interleaves:

from itertools import count

counter = count()

def increment():
    return next(counter)

The same three-step problem shows up in Java, where synchronized or an atomic wrapper handles the counter, and in Go, where the race detector catches it before it ships:

var count int64

func increment() {
    atomic.AddInt64(&count, 1)   // indivisible
}

File updates follow the same pattern. Two background jobs writing the same config file will not error, they will just produce a mixture of both versions:

# unsafe: read, edit, write as separate steps
sed -i "s/timeout=30/timeout=60/" config.ini

# safer: one atomic replace, so readers see old or new, never half
tmp=$(mktemp) && sed "s/timeout=30/timeout=60/" config.ini > "$tmp" && mv "$tmp" config.ini

Databases are where the money meets the bug. Two withdrawals reading the same balance and both writing back the result is a double-spend:

-- lost update: balance is read, then written back unchanged
SELECT balance FROM accounts WHERE id = 1;
UPDATE accounts SET balance = 50 WHERE id = 1;

-- atomic conditional update: the database does the read-modify-write
UPDATE accounts SET balance = balance - 100
WHERE id = 1 AND balance >= 100;

If the row affects zero rows, the withdrawal was refused. Nobody had to guess what the balance was in between.

What Is the Difference Between a Race Condition and a Data Race?

They sound interchangeable and they are not. A data race is a memory-model term: two threads access the same memory location, at least one access is a write, and the two accesses are not ordered by a happens-before relationship. A race condition is a logic term: the program’s result depends on timing, regardless of what the hardware did.

The gap matters. Some race conditions are harmless because the interleaving still produces a valid answer. Some data races are harmless because the racing writes land on the same value. And a genuine data race in C or C++ is undefined behaviour, which means the compiler is entitled to assume your program never has one and to optimise your code accordingly.

Here is the wider picture, including the two failure modes people confuse with races:

ProblemWhat happensTypical symptomHow you address it
Race conditionResult depends on operation orderingWrong count, lost update, double charge, intermittently bad dataAtomic operations, locks, transactions, ownership rules
Data raceUnsynchronised concurrent memory access, undefined behaviourGarbage values, torn reads, compiler-dependent nonsenseHappens-before edges: atomics, locks, thread joins, proper memory ordering
DeadlockThreads each hold a lock the other needsProgram hangs with no error and no progressConsistent lock ordering, timeouts, single lock where possible
StarvationA thread is perpetually denied a resourceOne task never finishes while the system stays busyFair scheduling, bounded priority, try-lock with backoff

One more term worth knowing is lost wakeup: one thread changes state and signals a waiter, but the other thread evaluated its wait condition before that change and has not re-checked yet. It goes back to sleep after the signal already fired. Every wait loop should re-check its condition inside a loop, never once on entry.

Why Are Race Conditions Hard to Reproduce?

Why Are Race Conditions Hard to Reproduce?

Because the trigger is timing, not state. The same input on the same machine can pass on Tuesday and fail on Wednesday, and nothing in your data changed.

Several things make it worse. The operating system scheduler decides when your thread runs, and that decision changes with load, core count and CPU frequency scaling. Different hardware changes the race window width. Compiler optimisations and CPU out-of-order execution move instructions around inside a single thread, which shifts the interleaving without changing your source.

Logging distorts what it is trying to record. Adding a log line inserts work into the timing window and can make the bug disappear, which is why a race that vanishes the moment you turn on verbose output is a race.

Debuggers do the same thing. Single-stepping serialises execution, so the interleaving that caused the failure never happens again. This is why races get dismissed as works on my machine.

AI-assisted review has the same blind spot, and the community has noticed. One developer-reported test found an AI code reviewer caught a race in a Go program roughly five to ten percent of the time, and the reaction in that thread was that human review is still the more reliable tool for concurrency bugs. If a model only spots it occasionally, you cannot use it as your safety net.

The practical consequence: you cannot test a race away. You can only build the code so the interleaving that matters cannot produce a wrong answer.

How Do You Avoid Race Conditions?

Prevention follows a short list, and the order matters because it goes from cheapest to most structural.

  1. Make the operation atomic so no interleaving can occur in the middle of it.
  2. Guard it with a lock around the narrowest region that needs protection.
  3. Wrap database changes in a transaction, using row locks, version columns or conditional updates.
  4. Give the state one owner and communicate through messages or queues instead of shared writes.
  5. Remove the check entirely when you can express the same requirement as a single operation.
  6. Make retries safe with idempotency keys so a repeated request cannot do the work twice.
  7. Test the concurrency deliberately with race detectors, barriers and repeated runs rather than trusting timing.

Use Locks and Synchronization Primitives

A mutex, or lock, allows one holder at a time inside a critical section. A semaphore limits how many threads may enter at once, which suits a bounded pool of resources. Condition variables let a thread wait for state it cannot poll safely on its own. Reader-writer locks let many readers share while giving writers exclusive access.

lock.acquire()
try:
    shared = read_state()
    shared.step_one()
    write_state(shared)
finally:
    lock.release()

Two rules matter more than the choice of primitive. Take the narrowest lock you can, because every extra instruction inside it is a place another thread waits. And always release on the failure path, which is why try/finally or a context manager beats a bare acquire.

Locks create their own failure mode. Two threads taking two locks in opposite order will deadlock, and the fix is a global ordering: every code path acquires locks in the same sequence, or a timeout that breaks the cycle. Consistent lock ordering plus a bounded wait is what turns a deadlock risk into a retryable error.

The trap most guides skip: an in-process lock does not span processes. In serverless, containers or anything with more than one instance, each instance has its own memory, so an application-level mutex protects one instance while the other carries on unprotected. Developers on r/devops described exactly this false confidence. When state has to be shared across instances, the lock has to live in shared infrastructure, or you redesign so there is nothing to lock.

Use Atomic Operations to Avoid a Race Condition on Shared State

For simple shared state, one indivisible instruction beats any amount of locking. Atomic increment adds one and returns the new value without any window between read and write. Compare-and-swap takes an expected value and only writes if the memory still holds it, which is how lock-free retry loops work.

// Java / C#
Interlocked.Increment(ref count);
Interlocked.CompareExchange(ref slot, expected, replacement);

// Go
atomic.AddInt64(&count, 1)
atomic.CompareAndSwapInt64(&slot, &expected, replacement)

// C / C++
atomic_fetch_add(&count, 1);
atomic_compare_exchange_strong(&slot, &expected, replacement);

Reach for atomics when the state is a single number, a flag, or a pointer you publish once. They are fast, they never deadlock, and they make the intent obvious to the next reader.

Two cautions. Atomicity does not give you memory visibility for free in C and C++, where you still need the right memory ordering for your acquire-release pattern. And an atomic is not a transaction: three separate atomic operations can still interleave with other threads between them, so composing several steps still needs a lock or a lock-free CAS loop.

Design Around Ownership and Immutability

The strongest fix is to remove the shared mutable state rather than defend it. If exactly one thread owns a piece of state and everything else sends it a message, there is no race to lose. This is how actor models, Erlang processes and most modern GUI architectures handle threading.

Message passing moves the concurrency question from memory to queues. A single-consumer queue guarantees ordering and removes contention entirely, which is also why a FIFO queue keeps coming up as the answer when several concurrent requests hit the same customer record.

Immutable values are the other half. If a shared object cannot change after construction, readers need no synchronisation at all. Instead of mutating a list in place, produce a new one and swap the reference atomically. Thread-safe collections such as ConcurrentHashMap or BlockingQueue apply the same principle inside a data structure you did not write.

Watch for hidden shared state inside libraries, singletons, module-level globals and connection pools. A race you cannot locate is usually living in a component nobody assumed was concurrent.

Make Database Updates Transactional

Databases give you the same guarantees as locks, but with durability and crash recovery attached. Wrap multi-statement changes in a transaction, and choose an isolation level that matches the requirement rather than the default.

-- pessimistic: hold the row for the duration
BEGIN;
SELECT balance FROM accounts WHERE id = 1 FOR UPDATE;
UPDATE accounts SET balance = balance - 100 WHERE id = 1;
COMMIT;
-- optimistic: version column, detect the conflict yourself
UPDATE accounts SET balance = balance - 100, version = version + 1
WHERE id = 1 AND version = 42;

Zero rows affected means someone else got there first, and you retry or return a conflict. That is optimistic concurrency control, and it suits low-contention workloads where holding a lock for the whole transaction would waste time.

For counters and single-row state, prefer one conditional statement over read-then-write. It is shorter, it is atomic by definition, and there is no window to lose.

Over HTTP, the remaining problem is duplicate delivery. Retries, client double-clicks and webhook redelivery all produce the same request twice, so make the handler idempotent: require an idempotency key, store it with the result, and return the stored response if the same key arrives again. Where your database supports advisory locks, they are also useful for serialising work across application instances.

Test Concurrency Instead of Relying on Timing

Testing a race needs intent. Random load finds some bugs and misses others silently, so pair stress testing with tooling that reasons about concurrency explicitly.

  1. Run a dynamic race detector. Go ships go test -race and go run -race; Java and C use ThreadSanitizer and AddressSanitizer; Rust and ThreadSanitizer-backed C toolchains report conflicting accesses directly. These insert instrumentation that flags unsynchronised access at runtime.
  2. Use barriers and latches in tests so every thread reaches the same point before any proceeds. A barrier turns an intermittent race into a reliable one.
  3. Control the interleaving directly. If you can pause a thread between the read and the write, you can test the exact failure deterministically instead of hoping to hit it.
  4. Assert invariants, not outputs. After the run, check that the counter equals iterations times threads, that no balance went negative, and that every queued job ran once.
  5. Run the test many times with varied core counts and under CPU pressure, since scheduling changes alter the window.
  6. Add static rules for check-then-act patterns. Semgrep and similar tools catch two-step file operations, permission checks followed by opens, and read-then-write on shared fields at review time.
  7. For web endpoints, send parallel requests. Tools built for HTTP request racing can hit a race window deliberately and tell you whether the response differs between attempts.

Detection tooling is the most-requested content on this topic, and the reason is simple: the bug will not reproduce on demand, so you need something that watches every access instead.

Race Conditions in Real Systems

Web servers and duplicate requests. A user double-clicks checkout, or a proxy retries a timed-out request. Two handlers read the same cart, both charge, and the customer sees one order and two payments. Idempotency keys on the request are what stop this.

Node.js and single-threaded-looking code. Every await is a gap where another request runs. Developers on r/node described validating a payment gateway redirect and then writing to the database: the await in between is long enough for a second request to pass the same check. The runtime having one thread does not make the program sequential, because your code suspends and resumes.

Shell scripts, cron jobs and automation. Two overlapping runs of the same cron job will read the same input file, both process it, and both write the output. The nastier variant is the symlink swap, where a script checks ownership with one call and reopens the path with another. Automation users on r/shortcuts hit the same overlap with their triggers firing again before the previous run finishes.

Two fixes work well here. A lock directory created atomically with mkdir makes a second run exit cleanly, and writing to a temporary path then renaming replaces the file in one step.

mkdir /var/lock/report.lock 2>/dev/null || { echo "already running"; exit 0; }
trap 'rmdir /var/lock/report.lock' EXIT

CI/CD pipelines. Two jobs deploying the same artifact, or a shared cache directory written in parallel, produce builds that differ depending on ordering. Give each job its own workspace and publish through a single atomic step.

Distributed systems. Distributed locks across containers look like the answer until a network partition splits the cluster: the lock service believes two owners hold the same lock, and the classic Redlock debate is largely about whether that is acceptable. Prefer designs that do not need a global lock, and use fencing tokens if you must take one.

Operating systems and home labs. Shared filesystems, concurrent VM operations and backup tooling all assume nothing else is modifying the same files. This is why snapshots during a running backup, or two hypervisors writing the same disk image, produce corruption rather than a clean error. Real CVEs have exploited this class directly, including CVE-2021-21315 and CVE-2020-24815, and the Git apply symlink race.

Frameworks rarely advertise their races either. Caching layers with lazy initialisation, connection pools and rate limiters are all shared mutable state wearing a friendly API.

Frequently Asked Questions

Can a race condition exist in single-threaded code?

Yes. A runtime with one execution thread still has concurrency whenever your code suspends. Every await, yield or callback boundary lets another request run in between, so a check-then-act pattern split across two awaits has the same flaw as a threaded version. Shell scripts have the same problem from a different direction: two separate processes running at once. The requirement is overlapping operations on shared state, not multiple threads.

Is every data race a race condition?

No, though they overlap heavily. A data race is the memory-model definition: two unsynchronised accesses to the same location where one writes, with no happens-before relationship. A race condition is the logical result: the outcome depends on timing. Some data races produce correct results by luck, and some logical races involve no unsynchronised memory at all, such as two requests reading the same database row. Fixing one does not automatically fix the other.

How do I debug an intermittent race condition?

Start by making it reproducible. Run a dynamic race detector such as go test -race or ThreadSanitizer, which flag conflicting accesses instead of waiting for a wrong answer. Add barriers in tests so all threads reach the same point before continuing, then assert the invariant rather than the output. Avoid single-stepping, since it serialises execution and hides the interleaving, and remember that added logging can close the race window too.

What is the safest way to update a shared counter?

Use the runtime’s atomic primitive where one exists: itertools.count in Python, AtomicInteger or Interlocked in Java and C#, or sync/atomic in Go. A single indivisible operation removes the read-modify-write window entirely and cannot deadlock. If the increment depends on other state, hold a lock for the whole calculation instead, and keep that region as small as you can so other threads are not blocked for longer than necessary.

Are race conditions the same as deadlocks?

No, and they need different fixes. A race condition produces a wrong result while the program keeps running. A deadlock stops progress: each thread holds a lock the other needs, and nothing times out by default. Adding a lock to fix a race can therefore introduce a deadlock, which is why consistent lock ordering and bounded wait times belong in the same review as the original fix.

Can race conditions be completely eliminated?

Not with discipline alone, but they can be made structurally impossible in the code you control. Give each piece of state a single owner, use atomic operations and transactions for what remains shared, and remove check-then-act patterns where one statement will do. You can still be surprised by a library that hides shared state or by a network partition you did not design for, so race detectors and concurrency tests stay part of the routine.

Conclusion

To avoid a race condition, start by finding your shared mutable state and writing down who can touch it. That single inventory removes most of the risk, because it turns a mysterious intermittent failure into a short list of places that need protecting.

Then work in this order: make single operations atomic, add locks only where an atomic cannot do the job, wrap database changes in transactions with a conditional update, and give any remaining shared state one owner that others message instead of writing to.

Finally, run a race detector in CI and add tests that use barriers and invariant checks. You will not be able to test a race away, but you can make the wrong interleaving impossible to express.

Leave a Comment