Silent Failures in Bash Scripts: set -e, pipefail and Where Strict Mode Abandons You
set -euo pipefail does not make your script safe, it only gets you started. -e switching off inside function bodies, pipefail failing on SIGPIPE, -u ignoring empty variables and quotes breaking in nested shells: the cause and the cure for each.
The most dangerous automation script is not the one that crashes. It is the one that keeps going. A step in the middle of a deploy fails, the remaining steps run anyway, the exit code is zero and the dashboard stays green. You find out weeks later, when the nightly backup turns out to have been empty since the day someone renamed a directory. This article is about why bash scripts fail so quietly and what to do about it. I assume you know the shell; the focus is on the traps that catch experienced people.
Three levels of rigor
There are three stages of script robustness, and most teams stop at the first.
Doing nothing. By default bash treats every command as independent: if one fails, it moves on to the next. That is the right behaviour in an interactive session and the wrong one in a script, especially when the failing command is cd and the next one is rm.
Strict mode. Starting every script with set -euo pipefail. -e exits on the first failing command, -u makes an unset variable an error, pipefail makes a pipeline fail if any stage fails. This is the best return on effort you will get, but it is a starting point, not a guarantee. As you will see below, all three have gaps, and unknown gaps make the confidence strict mode gives you more dangerous than no confidence at all.
Strict mode plus verification. Instead of asserting that the script did something, measuring the result: the file actually exists and has the expected size, the service actually reports healthy, the backup actually restores. Anything long-lived and important has to get here.
My position: every script starts at level two, and anything on a production path graduates to level three. What follows are the holes in level two.
set -e is off in more places than you think
Bash suspends the -e rule wherever a command's exit status is already being tested: if, while and until conditions, every element of an && or || chain except the last, and anything negated with !. So far, reasonable. The surprise is that the suspension extends into the entire body of a function called in one of those positions:
set -e
prepare() {
false # you expect the function to stop here
echo "still running" # it prints
}
if prepare; then echo "ready"; fi
Because prepare is called as an if condition, nothing inside it is protected by -e. The person calling the function has no way to see this and the person who wrote it usually does not know. The fix is for functions to handle their own errors explicitly: write the critical step as command || return 1 and never rely on -e inside a function.
The second hole is command substitution. path=$(compute) will trigger -e if compute fails, but local path=$(compute) will not: the exit status of that line belongs to local, and local always returns zero. The same applies to export. Separate the declaration from the assignment:
local path
path=$(compute)
On bash 4.4 and later, shopt -s inherit_errexit makes the subshell inside a command substitution inherit -e; by default it does not.
pipefail and consumers that leave early
pipefail is the right tool, but it produces false failures whenever a consumer in the pipeline exits before the producer is done writing. The classic case:
set -o pipefail
printf '%s\n' "$content" | grep -q -- "$pattern"
grep -q exits on the first match. If printf is still writing, it is now writing to a pipe with no reader, receives SIGPIPE and dies with status 141. pipefail makes 141 the status of the whole pipeline, your condition evaluates as false, and the script takes the "pattern not found" branch. You will never see it on small input; it appears once the content is multi-line and larger than the pipe buffer, such as the output of a command with thousands of lines, which means it appears in production.
The cure is to not need an early-exiting consumer. For small data, hold the content in a variable and use the shell's own matching: [[ $content == *"$pattern"* ]]. For large data, consume the whole output and count, or disable pipefail for that one line. The same trap wears other costumes: head -n1, an early exit inside awk, anything that writes to a pager.
set -u is only half a guard
-u stops on an unset variable. It says nothing about an empty one, and most catastrophic commands run with an empty variable:
rm -rf "$TARGET"/*
If TARGET is set to the empty string, -u is silent and the command sweeps the root filesystem. Before every destructive command, check separately that the variable is non-empty and has the shape you expect:
[[ -n $TARGET && $TARGET == /srv/* ]] || { echo "bad TARGET: '$TARGET'" >&2; exit 1; }
That line looks tedious. Because it looks tedious it does not get written, and because it does not get written there are incidents.
Warn-and-continue hides the failure
The pattern that quietly opts out of strict mode is command || echo "WARNING: step failed". The author believes they made a deliberate choice, but they skipped two questions: do the following steps still make sense if this one failed, and who is going to read the warning. The warning lands in a log line, the dashboard stays green, and the failure hides until its next victim.
The rule is simple. If a step is not optional, its failure fails the script. If it is optional, say so in the code and reflect the warning in the exit status, for example by exiting with 2 at the end when the warning count is above zero. Nobody's job is to grep for "WARNING" in the log of a script that exited zero.
Tab-separated fields and read
Tabs look like a convenient field separator for tabular data, and read will betray you for it. Whitespace characters in IFS (space, tab, newline) are collapsed when they occur in sequence. In a three-column row where the middle field is empty, two tabs sit side by side, read treats them as one separator, and the value of column three lands in column two. There is no error. The script simply processes the wrong record under the wrong label.
Setting IFS=$'\t' does not help; tab is still a whitespace character. For data that can carry empty fields, pick a non-whitespace separator such as |, or hand the parsing to a tool that preserves field count, such as awk -F'\t'. Better still, if the data is structured, emit JSON and read it with jq.
Carrying quotes into a remote shell
In nested shell layers such as ssh host "..." or docker exec ctr bash -c '...', a comment, and in particular a word with an apostrophe in it, closes the single-quoted string early. Whatever follows is interpreted by the outer shell, the variables there expand to nothing, and the rm scenario above runs on your own machine. The rule: never put comments inside a nested shell string, and send anything longer than a line or two as a heredoc instead of escaping quotes:
ssh host 'bash -s' <<'EOF'
set -euo pipefail
cd /var/lib/app
./update.sh
EOF
Quoting the delimiter (<<'EOF') stops the local shell from expanding the content; the script arrives exactly as written.
Five habits
Put set -euo pipefail at the top, knowing the holes above. Add an ERR trap so you know where a failure happened; trap 'echo "error at line $LINENO" >&2' ERR alone saves hours. Create temporary files with mktemp and remove them in an EXIT trap. Wrap every external call that could hang in timeout; a script waiting forever is also a silent failure. Run every script through shellcheck, which catches unquoted variables, untested cd and misuse of read before you ever execute it.
Then sabotage your own guards. Invert the condition, empty the variable, and watch the script turn red with your own eyes. A guard that stays green under sabotage is not guarding anything.
When not to write bash
If the script is past a hundred lines, produces or parses JSON or YAML, needs retries or concurrency, or will be maintained by more than one person, bash is the wrong tool. The shell is excellent at connecting commands and poor at carrying logic. Once you cross that line, moving to a language with error handling built in, Python or Go, removes every trap in this article in one step. The moment you add the third nested if to a bash script is the moment to move.