Web scraping is one of those projects that looks simple at first:
«Send a request → get a webpage → extract some information.»
But once you start implementing it, you quickly discover that there are several pieces involved.
Recently, I started learning how to build a web scraper in Go. What interested me most was that I could build most of the foundation using Go's standard library.
In this article, I'll break down the main packages involved, what each one does, and how they fit together to create a simple web scraper.
- The basic architecture
Before looking at individual packages, it helps to understand the flow.
A basic scraper can be thought of as:
URL
│
▼
HTTP Request
│
▼
HTTP Response
│
▼
Response Body
│
▼
HTML/Text Processing
│
▼
Extract Information
│
▼
Validate / Store Data
In Go, different standard library packages can handle different parts of this process.
Some of the important ones are:
net/http
io
net/url
strings
net/http/httptest
Let's look at each one.
- "net/http" — communicating with websites
The first thing a scraper needs to do is communicate with a web server.
This is where the "net/http" package comes in.
A simple HTTP GET request looks like this:
package main
import (
"fmt"
"net/http"
)
func main() {
resp, err := http.Get("https://example.com")
if err != nil {
fmt.Println(err)
return
}
defer resp.Body.Close()
fmt.Println("Status:", resp.Status)
}
The important part is:
resp, err := http.Get("https://example.com")
This sends an HTTP GET request to the server.
The server responds with an "http.Response".
The response contains information such as:
resp.Status
resp.StatusCode
resp.Header
resp.Body
For example:
fmt.Println(resp.StatusCode)
might produce:
200
A "200" response generally means the request was successful.
- Why "defer resp.Body.Close()" matters
You'll often see this immediately after making an HTTP request:
defer resp.Body.Close()
Why?
Because the response body is a resource that needs to be closed when we're finished with it.
Think of it like opening a file.
You open it:
Open
↓
Read
↓
Close
HTTP response bodies follow a similar pattern.
So this:
resp, err := http.Get(url)
if err != nil {
return
}
defer resp.Body.Close()
means:
«Once this function finishes, close the response body.»
This is important for avoiding resource leaks.
- "io" — reading the response
Getting the response isn't enough.
The actual webpage content is inside:
resp.Body
But "resp.Body" is an "io.ReadCloser".
We need to read it.
One simple approach is:
body, err := io.ReadAll(resp.Body)
if err != nil {
fmt.Println(err)
return
}
fmt.Println(string(body))
The "io.ReadAll()" function reads everything from the reader.
The result is a byte slice:
[]byte
That's why we convert it to a string:
string(body)
For example:
fmt.Println(string(body))
could produce HTML like:
<!DOCTYPE html>
Example Domain
Now we have the webpage content in memory.
- Understanding "io.Reader"
One of the concepts I found particularly useful while learning this was the idea of an "io.Reader".
Go uses interfaces extensively.
Many things can be represented as a reader:
File
│
├── Reader
│
HTTP Response Body
│
├── Reader
│
Memory buffer
│
└── Reader
This means functions such as:
io.ReadAll()
can work with many different sources.
For example:
data, err := io.ReadAll(resp.Body)
The function doesn't need to know that the data came from a website.
It only needs something that implements the "io.Reader" interface.
This is one of the ideas that makes Go's standard library very composable.
- "net/url" — working with URLs
When building a real scraper, URLs quickly become more complicated.
You may need to:
- parse URLs
- inspect their components
- modify query parameters
- construct URLs
- resolve relative links
Go provides the "net/url" package for this.
For example:
package main
import (
"fmt"
"net/url"
)
func main() {
rawURL := "https://example.com/products?page=2"
u, err := url.Parse(rawURL)
if err != nil {
fmt.Println(err)
return
}
fmt.Println("Scheme:", u.Scheme)
fmt.Println("Host:", u.Host)
fmt.Println("Path:", u.Path)
fmt.Println("Query:", u.RawQuery)
}
The output would be something like:
Scheme: https
Host: example.com
Path: /products
Query: page=2
Instead of treating a URL as just a string, we can work with its individual components.
- Query parameters
Suppose our scraper needs to request:
https://example.com/products?page=2
Instead of manually constructing the string, we can use "url.Values".
params := url.Values{}
params.Set("page", "2")
u := url.URL{
Scheme: "https",
Host: "example.com",
Path: "/products",
RawQuery: params.Encode(),
}
fmt.Println(u.String())
This produces:
https://example.com/products?page=2
This becomes especially useful when building scrapers that need pagination.
For example:
/products?page=1
/products?page=2
/products?page=3
/products?page=4
- "strings" — processing extracted data
Once we've downloaded a webpage, we'll often need to work with text.
The "strings" package provides many useful functions.
For example:
title := " Example Product "
title = strings.TrimSpace(title)
fmt.Println(title)
Output:
Example Product
We can also search for text:
if strings.Contains(title, "Product") {
fmt.Println("Found product")
}
Or split strings:
links := "home,products,contact"
parts := strings.Split(links, ",")
for _, link := range parts {
fmt.Println(link)
}
Output:
home
products
contact
These operations become useful after extracting information from HTML.
- A simple scraper using these packages
Now let's combine what we've learned.
Suppose we want to download a webpage and print its HTML.
package main
import (
"fmt"
"io"
"net/http"
)
func main() {
url := "https://example.com"
resp, err := http.Get(url)
if err != nil {
fmt.Println("Request failed:", err)
return
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
fmt.Println("Unexpected status:", resp.Status)
return
}
body, err := io.ReadAll(resp.Body)
if err != nil {
fmt.Println("Failed to read response:", err)
return
}
fmt.Println(string(body))
}
The flow is now:
http.Get()
↓
http.Response
↓
resp.Body
↓
io.ReadAll()
↓
[]byte
↓
string
That's the foundation of our scraper.
- Extracting information from the response
Downloading the HTML is only half the job.
A scraper is useful because it extracts specific information.
For example, imagine the response contains:
Go Web Scraping
Learning Go is interesting.
We could perform very simple text processing with "strings".
html := `
Go Web Scraping
Learning Go is interesting.
`
if strings.Contains(html, "Go Web Scraping") {
fmt.Println("Title found")
}
However, this approach becomes fragile very quickly.
HTML is structured data, so for serious HTML parsing, a dedicated HTML parser is usually more appropriate.
For example, Go developers commonly use:
golang.org/x/net/html
This is not part of the Go standard library, but it works naturally with the standard library's HTTP and I/O packages.
- Parsing HTML with "golang.org/x/net/html"
After downloading a webpage, we can parse the HTML into a document tree.
Conceptually:
HTML Document
│
▼
│
┌────┴─────┐
▼ ▼
│
┌────┴────┐
▼ ▼
The parser allows us to navigate this structure instead of treating HTML as plain text.
A simplified example:
doc, err := html.Parse(resp.Body)
if err != nil {
fmt.Println(err)
return
}
Now "doc" represents the parsed HTML document.
From there, we can walk through the nodes and look for elements such as:
- Finding links
Links are particularly important when building a crawler.
Suppose we have:
A crawler could find those "href" attributes and then visit them.
Conceptually:
Page A
│
├── /about
│
└── /products
│
▼
Page B
This is where web scraping starts becoming web crawling.
A scraper might extract information from one page.
A crawler can follow links and discover additional pages.
- "net/http/httptest" — testing the scraper
One of the most useful packages I encountered was:
net/http/httptest
Testing a scraper against a real website is not always a good idea.
The website could:
- change its HTML
- become unavailable
- block requests
- return different data
- become slow
Instead, we can create a fake HTTP server.
Example:
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
fmt.Fprintln(w, "
Hello
")}))
defer server.Close()
Now we have a local test server.
We can send our scraper to:
server.URL
instead of a real website.
- Testing a scraper with a fake server
For example:
func TestScraper(t *testing.T) {
server := httptest.NewServer(http.HandlerFunc(
func(w http.ResponseWriter, r *http.Request) {
fmt.Fprintln(w, "
Hello
")},
))
defer server.Close()
resp, err := http.Get(server.URL)
if err != nil {
t.Fatal(err)
}
defer resp.Body.Close()
body, err := io.ReadAll(resp.Body)
if err != nil {
t.Fatal(err)
}
if !strings.Contains(string(body), "Hello") {
t.Fatal("expected Hello in response")
}
}
Now our test doesn't depend on the internet.
We control exactly what the server returns.
That's extremely useful.
- Testing different HTTP responses
We can also test how our scraper handles errors.
For example:
server := httptest.NewServer(http.HandlerFunc(
func(w http.ResponseWriter, r *http.Request) {
http.Error(w, "Not Found", http.StatusNotFound)
},
))
Now our scraper receives:
404 Not Found
We can verify that our program handles that situation correctly.
We can also test:
200 OK
301 Redirect
403 Forbidden
404 Not Found
500 Internal Server Error
This is where testing becomes much more valuable than simply checking whether the scraper works once.
- Putting the packages together
At this point, the roles of the packages become clearer.
Package| Purpose
"net/http"| Send HTTP requests and receive responses
"io"| Read response data
"net/url"| Parse and construct URLs
"strings"| Process textual data
"net/http/httptest"| Test HTTP-based applications
"golang.org/x/net/html"| Parse HTML documents
They aren't competing packages.
They form a pipeline.
URL
│
▼
net/url
│
▼
net/http
│
▼
HTTP Response
│
▼
io
│
▼
HTML data
│
▼
x/net/html
│
▼
Extract data
│
▼
strings
- What I learned beyond the packages
The biggest lesson for me wasn't actually the individual packages.
It was understanding how they work together.
Before this, it was easy to think of web scraping as:
Download webpage
Extract information
But a more realistic system looks like:
URL
↓
HTTP request
↓
Timeout handling
↓
HTTP status validation
↓
Response body
↓
HTML parsing
↓
Data extraction
↓
Data validation
↓
Storage
↓
Tests
And once you start building a real scraper, additional concerns appear:
- retries
- timeouts
- redirects
- rate limiting
- robots.txt
- concurrent requests
- duplicate URLs
- malformed HTML
- authentication
- logging
- error handling
- data persistence
That's where a simple scraper starts turning into a real software engineering project.
- A small mental model
The way I'm currently thinking about Go web scraping is:
net/http
↓
"Talk to the website"
io
↓
"Read what the website sent"
net/url
↓
"Understand the address"
x/net/html
↓
"Understand the HTML structure"
strings
↓
"Clean and process text"
httptest
↓
"Test everything without depending on the real website"
Once you understand those responsibilities, the code becomes much easier to reason about.
- What's next?
My next step is to move beyond a basic scraper and explore how to build a more robust crawler in Go.
Some areas I want to explore are:
- HTTP clients and custom timeouts
- HTML tree traversal
- Extracting links and images
- URL normalization
- Concurrent scraping with goroutines
- Rate limiting
- Retry strategies
- Structured data models
- Testing edge cases
- Persisting scraped data
- Handling failed requests
- Building a complete CLI scraper
The interesting part is that the deeper I go, the more I realize that web scraping is not really about scraping.
It's about understanding how different parts of a software system communicate.
And Go's standard library provides many of the building blocks needed to understand and build those systems.
Top comments (0)