What Is a Regular Expression? A Complete Explanation
Quick Answer
A regular expression is a pattern that describes a set of strings. It uses literal characters (match themselves) and metacharacters (*, +, ?, ., ^, $, [], (), {}, |, \) with special meanings. Regex is used for validation, searching, extraction, and replacement. Use our free Regex Tester to test patterns against text in your browser.
Introduction
A regular expression (regex or regexp) is a sequence of characters that defines a search pattern. The pattern is used by string-searching algorithms to find, match, or replace text. Regular expressions originated in formal language theory in the 1950s (Stephen Kleene) and were adopted into Unix text tools (grep, sed, awk) in the 1970s. Today, regex is built into virtually every programming language and text editor. Regex is powerful — it can validate email formats, extract data from logs, and find complex patterns — but the syntax is dense and notoriously hard to read.
Step by Step
-
Understand literal and metacharacters
Literal characters (a, b, 1, @) match themselves. Metacharacters have special meanings: . matches any character, * means zero or more, + means one or more, ? means zero or one, ^ matches start, $ matches end.
-
Learn character classes
Square brackets [] define a character class — [aeiou] matches any vowel, [0-9] matches any digit, [^0-9] matches any non-digit. Shorthand classes: \d (digit), \w (word character), \s (whitespace), and their uppercase negations.
-
Use quantifiers
Quantifiers specify how many times a pattern repeats: * (0+), + (1+), ? (0 or 1), {n} (exactly n), {n,} (n+), {n,m} (n to m). Greedy by default; add ? for lazy matching (e.g., *?).
-
Group and anchor
Parentheses () create capture groups. Alternation | matches either side (cat|dog). Anchors ^ and $ match start and end of string. Use these to build complex patterns from simpler ones.
Examples
Match an email-like pattern
Input: Pattern: [a-z]+@[a-z]+\.[a-z]+
Output: Matches: alice@example.com, bob@test.org
Match a phone number format
Input: Pattern: \d{3}-\d{3}-\d{4}
Output: Matches: 555-123-4567
Extract key-value pairs
Input: Pattern: (\w+)=(\w+)
Output: Matches 'key=value' with group 1 as key, group 2 as value
Common Problems
- Greedy matching over-consumes —.* matches as much as possible. Use .*? for lazy matching to find the shortest match.
- Forgetting to escape special characters —., *, +, ?, etc. are special. Use \ to match them literally (e.g., \. matches a literal dot).
- Catastrophic backtracking —patterns with nested quantifiers (e.g., (a+)+) can take exponential time on certain inputs, causing ReDoS attacks.
- Over-matching —without anchors (^ and $), a pattern matches anywhere. Use ^...$ to match the entire string.
Tips
- Start simple and add complexity incrementally —debugging a complex regex all at once is very hard.
- Use a regex tester to experiment with patterns against sample text —see matches highlighted in real time.
- Comment complex regexes using extended mode (/x flag in some engines) or break them into named parts.
- Use our Regex Tester to test patterns instantly in your browser —no signup, fully private.