Building practical shell pipelines

Advanced linux users often rely on the terminal and complex-looking command chains to parse command output or log files. While it looks difficult, the vast majority of it is made up of a small set of tools.

What to expect

This guide will introduce you to the most important data processing standard unix programs and their most common parameters. We assume you already know that pipes | redirect the output of one command to the input of the next, and have basic linux terminal knowledge.


This guide is by no means an exhaustive explanation of command-line scripting or a collection of one-liners, but rather meant to teach you how to build your own one-liners on the fly instead of just memorizing them.

grep,head and tail

The most basic need on the command line is restricting an input stream to a subset of lines you are actually interested in. The vast majority of this requirement can be satisfied with just three core unix tools:

  • head prints the first few lines only, specify how many with -n number, for example head -n 40 to return the first 40 lines

  • tail does the same as head, but instead working from the end of the input, meaning it returns the last n lines instead of the first

  • grep is used to return only lines matching a regular expression. Get used to always supplying -E for a more common regular expression syntax

With just these three tools, you can easily filter a log file sample.log to find all lines in the last 100 entries beginning with "error":

tail -n 100 sample.log | grep -E '^[Ee]rror'

Navigating log files largely consists of just these commands.

sort and uniq

Once you need more specific information from a log file, you may want to filter out duplicate lines to see how many different error messages there are:

sort sample.log | uniq

Using sort first sorts the input lines to that duplicates are consecutive - this is important, because uniq can only find unique lines if they are already sorted. When passing the -c flag to uniq, it will prefix each line with the number of times it existed in the original input.


All of these tools can be combined for one of the standard one-liner patterns you will see everywhere:

sort sample.log | uniq -c | sort -nr | head -n 10

The command first sorts the lines in sample.log, then strips duplicate lines and prefixes each unique line with the number of times it was found. Since all lines now start with that number, the second sort uses -n to sort them numerically and -r to revert the sorting order, displaying the most frequently found lines at the top. Finally, head -n 10 restricts the output to only 10 lines.


The pattern basically answers the question "which are the top 10 most frequently found lines in sample.log, sorted by number of occurrences?".

sed

The sed command is a stream editor, allowing string transformation on the fly. While it can do a lot, terminal usage is typically limited to one of three use cases: Replacing the first instance of a pattern, replacing all instances of a pattern, or deleting a line matching a pattern.


Replacing the first instance of a regular expression is simple:

sed -E 's/oldpattern/newstring/'

Starting with s means "substitution", aka "replacement" mode. It must be followed by a delimiter that separates the regular expression(s). By convention you would use / as the delimiter since it is not part of regex syntax and rarely used in patterns. If you do need it inside patterns, a good fallback is #. Remember to always use -E for common regular expression support.


Both oldpattern and newstring support regular expressions, including capturing groups and backreferences. So you could replace a line with a subset of its content:

sed -E 's/CPU Usage: ([0-9]+)%/\1/'

A line "CPU Usage: 13%" would get replaced by just "13" because it captures the usage number with a () capturing group and uses whatever was captured there to replace the line (\1 pastes whatever the first capture group matched).


To replace all occurrences of a pattern with something instead of just the first per line, add the g modifier after the patterns:

sed -E 's/oldpattern/newstring/g'

The last common function is deleting a line that matches a pattern:

sed -E '/pattern/d'

The d modifier deletes the entire line if the regular expression defined by pattern matches it.


Most often sed is used for normalization before processing with other tools in a pipeline. For example, you could collapse all consecutive whitespace characters into a single space:

sed -E 's/[[:space:]]/ /g'

Or merge all error-adjacent log messages into an easier to count format:

sed -E 's/^(error|warn|fail|fatal)/error/'

Or perhaps exclude all debug lines from a log file:

sed -E '/[Dd]ebug/d'

All of these examples are very basic and require much more regular expression knowledge than scripting or sed specifics, but one-liners will very rarely go beyond such simple usage.

awk

Easily the most complex and capable tool in this list, common usage boils down to a very limited subset.


The awk command automatically splits every input line into "fields", using whitespace as a delimiter. You can then call print $fieldNumber to replace the entire line with it, or $0 to get the original line.

awk '{print $3}' sample.log

You can also display multiple field numbers in any order you like, even joined with custom strings inbetween:

awk '{print "Third: " $3 ", Fourth: " $4}' sample.log

The print call automatically joins consecutive strings and field names into a single output line.

By default, fields are split by whitespace (can be more than one), but if you want another separator like : to parse /etc/passwd, you can change it with the -F flag:

awk -F':' '{print $7}' /etc/passwd

The example splits every line in /etc/passwd by : and prints the 7th field, which happens to be the login shell for each user.


The true power of awk comes from combining field extraction with conditional logic, similar to if conditions in other languages:

awk '$3 > 80 {print $0}' sample.log

While very short, the command excludes all lines where the third field has a value smaller than 80. String to number conversion happens automatically, and the common operators >, <, == and != are available. Additionally, a conditional can also match agaisnt a regular expression:

awk '$2 ~ /^[Ee]rror/ {print $0}' sample.log

Now only lines beginning with "Error" or "error" are kept.


Lastly, awk supports aggregate logic by appending a second executable block {} after the END keyword:

awk '{sum += $4} END {print sum}' sample.log

For every line, the value of the 4th field is added to the temporary sum variable. After all lines have been processed, the program after END is executed, printing the total sum of all field values.


The awk is basically a complete programming language, but command-line usage boils down to just those few common features.

Building pipelines

Once you have internalized just these few basic commands and flags, you can combine them to answer complex questions from raw data on the fly.


Let's assume you have an access.log file from an apache2 webserver:

192.168.1.10 - - [25/Aug/2026:19:30:01 +0200] "GET / HTTP/1.1" 200 4521
192.168.1.11 - - [25/Aug/2026:19:30:04 +0200] "GET /index.html" 200 8234
192.168.1.12 - - [25/Aug/2026:19:30:09 +0200] "POST /login HTTP/1.1" 302 512
192.168.1.10 - - [25/Aug/2026:19:30:15 +0200] "GET /css/style.css" 200 19342
192.168.1.13 - - [25/Aug/2026:19:30:21 +0200] "GET /api/users HTTP/1.1" 200 1208
192.168.1.14 - - [25/Aug/2026:19:30:27 +0200] "GET /images/logo.png" 200 34892
192.168.1.15 - - [25/Aug/2026:19:30:33 +0200] "GET /missing HTTP/1.1" 404 721
192.168.1.11 - - [25/Aug/2026:19:30:39 +0200] "GET /dashboard HTTP/1.1" 200 6842
192.168.1.16 - - [25/Aug/2026:19:30:45 +0200] "POST /api/logout HTTP/1.1" 204 0
192.168.1.17 - - [25/Aug/2026:19:30:52 +0200] "GET /robots.txt HTTP/1.1" 200 68

This is a brief sample, real log files contain thousands to millions of lines.


Lets say we want to know what HTTP response codes are most frequently sent, with counts.

There are a few tasks involved to solve this problem. First of all, we cannot split fields by whitespace directly since the double quotes can contain two or three entries, so that has to be normalized or removed. Afterwards, we can apply the correct field and apply the common uniq pattern we saw earlier:

sed -E 's/"[^"]+"//' access.log | awk '{print $6}' | sort | uniq -c | sort -nr

Note that we extract the field $6 because the entire quoted string including the quotes is removed by sed.


The output would look like:

7 200
1 404
1 302
1 204

This is how real world pipelines are built on the fly. Seasoned linux administrators do not memorize complete one-liners, but rather smaller commands like these and stitch them together as needed.

About cut and tr

You may run across the commands tr and cut in terminal command chains. They can be used, but their most common use cases can be emulated by other tools already discussed previously.


The tr command is used to translate or modify input character by character, most commonly with something like tr -s ' ' to collapse consecutive whitespace into a single space for easier field separation. It has other uses like stripping all symbols not in a charset or transforming lowercase to uppercase, but those are mostly niche in common command line pipelines. Tasks like normalizing whitespace or characters can be done using sed regular expression replacements as well, largely making tr redundant as a tool to build command chains on the fly.


The cut command's main use was to extract fields from lines of input. However, it requires a single separator, meaning multiple whitespace characters have to be collapsed first. The awk command can handle this much easier without preprocessing, and additionally brings conditionals and even access to the original line independent of conditional logic.

It is still useful for cutting sequences of characters or bytes out of input lines, but that is a fairly rare requirement for most tasks.

More articles

Mounting filesystems on linux

Including device identifiers, mount options and common errors

Essential incus operator guide

The 80% you actually need for most use cases