Extract URLs from Text
Extract all HTTP, HTTPS, FTP, and mailto links from unstructured text, HTML source code, Markdown manuscripts, or server access logs with instant deduplication and tracking UTM cleaning.
Source Text Input
Extracted Links Output
0 URLsParsed Link Directory
Individual Link Actions| # | Scheme | Domain / Host | Path & Query | Action |
|---|---|---|---|---|
| Paste text above to view structured links table. | ||||
How to Extract URLs from Unstructured Text
To extract web links: 1. Scan text with an RFC 3986 URI regular expression matching protocol schemes (https://, http://, ftp://, mailto:) or www. 2. Trim trailing sentence punctuation (. , ; : ! ? ) ] } " ') that adheres to natural language syntax. 3. Optionally normalize query parameters by removing marketing tracking tags (utm_*, gclid, fbclid). 4. Deduplicate unique URLs and sort by occurrence, alphabet, or root domain.
Uniform Resource Identifier (URI) Architecture & RFC 3986
In computer networking and web architecture, a Uniform Resource Identifier (URI) is a compact sequence of characters that identifies an abstract or physical resource. Defined by the Internet Engineering Task Force (IETF) in RFC 3986, every generic URI is composed of five hierarchical components:
Indicates the network protocol utilized to access the resource, such as https (Hypertext Transfer Protocol Secure), http, ftp, or mailto.
Encompasses the fully qualified domain name (FQDN) or IP address, optional user credentials, and custom network ports (e.g. example.com:8080).
Specifies the specific directory or endpoint resource on the server, organized in slash-separated segments (e.g. /web-tools/extract-urls/).
The query string begins with ? containing key-value parameters. The fragment begins with # pointing to an internal document anchor.
Regular Expression Mechanics & Boundary Tokenization
Extracting hyperlinks from raw text is deceptively challenging because URLs embedded in natural language prose frequently collide with sentence punctuation. For instance, in the sentence “You can verify the dataset at https://example.com/data.”, a naive whitespace tokenizer erroneously includes the trailing period as part of the path.
/(?:(?:https?|ftps?):\/\/|mailto:|www\.)[^\s<>"'`{}|\\^]+[^\s<>"'`{}|\\^.,;:!?)\\]}]/gi This regular expression requires valid protocol schemes or schema-less prefixes, matches non-whitespace character runs, and enforces that the terminating character excludes common punctuation marks like commas, periods, colons, closing parentheses, and angle brackets.
Critical Punctuation Trimming Rules
- Parenthetical Citations: When authors write
(see https://example.com/paper), the trailing closing parenthesis must be severed from the URL unless an opening parenthesis exists inside the URL query itself. - Terminal Sentence Full Stops: Any period (
.) occurring at the final character position is classified as grammar punctuation and trimmed. - HTML Attribute Quotes: Quotes enclosing attributes in
<a href="https://example.com">are excluded.
URL Hygiene: Deduplication, Tracking UTM Stripping & Normalization
Raw link dumps from emails, social media feeds, and advertising campaigns often contain redundant tracking parameters that obscure canonical targets. Cleaning these parameters provides true URL hygiene:
- • utm_source, utm_medium, utm_campaign (Google Analytics)
- • gclid, gclsrc (Google Ads Click IDs)
- • fbclid (Meta / Facebook Click Identifiers)
- • mc_eid, mc_cid (Mailchimp Campaign Tokens)
- • msclkid (Microsoft Advertising)
Stripping tracking parameters allows genuine deduplication. For example, example.com/?utm_source=fb and example.com/?utm_source=twitter resolve to the identical canonical landing page, preventing inflated link counts during migrations.
Practical Applications: SEO Audits, Security Incident Response & DevOps
Automated link extraction is a core operational capability across modern technical disciplines:
Catalog every external and internal link in a blog post, verify anchor text consistency, identify non-HTTPS mixed content links, and batch check for 404 broken URLs.
Extract URLs from phishing email headers, incident tickets, and threat intelligence feeds to feed into firewall blacklists, domain reputation checkers, and sandboxes.
Parse Nginx, Apache, or Kubernetes ingress logs to extract upstream API targets, verify outbound webhook configurations, and audit external cloud asset dependencies.
Step-by-Step Worked Extraction Examples
Review these real-world sample inputs showing how the engine tokenizes diverse formats:
Comprehensive URL Anatomy & Structural Comparison Table
| Component | RFC 3986 Syntax | Sample Value | Extraction Handling |
|---|---|---|---|
| Scheme | ALPHA *( ALPHA / DIGIT / “+” / “-” / “.” ) | https://, ftp:// | Mandatory prefix anchor (or www.) |
| Host / FQDN | IP-literal / IPv4address / reg-name | basicmathtools.com | Case-insensitive; used for domain grouping |
| Port | *DIGIT (0 – 65535) | :8443, :8080 | Preserved with authority when present |
| Path | path-abempty / path-absolute | /web-tools/extract-urls/ | Case-sensitive; trailing slashes retained |
| Query | *( pchar / “/” / “?” ) | ?page=2&sort=desc | Preserved; tracking UTMs optionally stripped |
| Fragment | *( pchar / “/” / “?” ) | #faqs, #top | Optional anchor; can be stripped via toggle |
Common Pitfalls & False-Positive Boundary Traps
Relative links like /about/ or ../images/logo.png lack an authority domain and cannot be reliably extracted without knowing the base origin URL of the hosting document.
In Markdown syntax like [Google](https://google.com), extracting without boundary awareness appends the closing bracket ) to the URL, resulting in a broken 404 target.
Lead Developer & Founder of Basic Math Tools. Specializes in browser-native computational algorithms and applied mathematics.
Mathematics & curriculum specialists. Audited against standard algebraic and arithmetic principles.