DEV Community

Ouma Asoyoh
Ouma Asoyoh

Posted on

Building a web scraper in go: Understanding the go standard libraries.

Web scraping is one of those projects that looks simple at first:

«Send a request → get a webpage → extract some information.»

But once you start implementing it, you quickly discover that there are several pieces involved.

Recently, I started learning how to build a web scraper in Go. What interested me most was that I could build most of the foundation using Go's standard library.

In this article, I'll break down the main packages involved, what each one does, and how they fit together to create a simple web scraper.


  1. The basic architecture

Before looking at individual packages, it helps to understand the flow.

A basic scraper can be thought of as:

URL
│
▼
HTTP Request
│
▼
HTTP Response
│
▼
Response Body
│
▼
HTML/Text Processing
│
▼
Extract Information
│
▼
Validate / Store Data

In Go, different standard library packages can handle different parts of this process.

Some of the important ones are:

net/http
io
net/url
strings
net/http/httptest

Let's look at each one.


  1. "net/http" — communicating with websites

The first thing a scraper needs to do is communicate with a web server.

This is where the "net/http" package comes in.

A simple HTTP GET request looks like this:

package main

import (
"fmt"
"net/http"
)

func main() {
resp, err := http.Get("https://example.com")
if err != nil {
fmt.Println(err)
return
}

defer resp.Body.Close()

fmt.Println("Status:", resp.Status)
Enter fullscreen mode Exit fullscreen mode

}

The important part is:

resp, err := http.Get("https://example.com")

This sends an HTTP GET request to the server.

The server responds with an "http.Response".

The response contains information such as:

resp.Status
resp.StatusCode
resp.Header
resp.Body

For example:

fmt.Println(resp.StatusCode)

might produce:

200

A "200" response generally means the request was successful.


  1. Why "defer resp.Body.Close()" matters

You'll often see this immediately after making an HTTP request:

defer resp.Body.Close()

Why?

Because the response body is a resource that needs to be closed when we're finished with it.

Think of it like opening a file.

You open it:

Open
↓
Read
↓
Close

HTTP response bodies follow a similar pattern.

So this:

resp, err := http.Get(url)

if err != nil {
return
}

defer resp.Body.Close()

means:

«Once this function finishes, close the response body.»

This is important for avoiding resource leaks.


  1. "io" — reading the response

Getting the response isn't enough.

The actual webpage content is inside:

resp.Body

But "resp.Body" is an "io.ReadCloser".

We need to read it.

One simple approach is:

body, err := io.ReadAll(resp.Body)

if err != nil {
fmt.Println(err)
return
}

fmt.Println(string(body))

The "io.ReadAll()" function reads everything from the reader.

The result is a byte slice:

[]byte

That's why we convert it to a string:

string(body)

For example:

fmt.Println(string(body))

could produce HTML like:

<!DOCTYPE html>


Example Domain


Example Domain



Now we have the webpage content in memory.


  1. Understanding "io.Reader"

One of the concepts I found particularly useful while learning this was the idea of an "io.Reader".

Go uses interfaces extensively.

Many things can be represented as a reader:

File
│
├── Reader
│
HTTP Response Body
│
├── Reader
│
Memory buffer
│
└── Reader

This means functions such as:

io.ReadAll()

can work with many different sources.

For example:

data, err := io.ReadAll(resp.Body)

The function doesn't need to know that the data came from a website.

It only needs something that implements the "io.Reader" interface.

This is one of the ideas that makes Go's standard library very composable.


  1. "net/url" — working with URLs

When building a real scraper, URLs quickly become more complicated.

You may need to:

  • parse URLs
  • inspect their components
  • modify query parameters
  • construct URLs
  • resolve relative links

Go provides the "net/url" package for this.

For example:

package main

import (
"fmt"
"net/url"
)

func main() {
rawURL := "https://example.com/products?page=2"

u, err := url.Parse(rawURL)
if err != nil {
    fmt.Println(err)
    return
}

fmt.Println("Scheme:", u.Scheme)
fmt.Println("Host:", u.Host)
fmt.Println("Path:", u.Path)
fmt.Println("Query:", u.RawQuery)
Enter fullscreen mode Exit fullscreen mode

}

The output would be something like:

Scheme: https
Host: example.com
Path: /products
Query: page=2

Instead of treating a URL as just a string, we can work with its individual components.


  1. Query parameters

Suppose our scraper needs to request:

https://example.com/products?page=2

Instead of manually constructing the string, we can use "url.Values".

params := url.Values{}

params.Set("page", "2")

u := url.URL{
Scheme: "https",
Host: "example.com",
Path: "/products",
RawQuery: params.Encode(),
}

fmt.Println(u.String())

This produces:

https://example.com/products?page=2

This becomes especially useful when building scrapers that need pagination.

For example:

/products?page=1
/products?page=2
/products?page=3
/products?page=4


  1. "strings" — processing extracted data

Once we've downloaded a webpage, we'll often need to work with text.

The "strings" package provides many useful functions.

For example:

title := " Example Product "

title = strings.TrimSpace(title)

fmt.Println(title)

Output:

Example Product

We can also search for text:

if strings.Contains(title, "Product") {
fmt.Println("Found product")
}

Or split strings:

links := "home,products,contact"

parts := strings.Split(links, ",")

for _, link := range parts {
fmt.Println(link)
}

Output:

home
products
contact

These operations become useful after extracting information from HTML.


  1. A simple scraper using these packages

Now let's combine what we've learned.

Suppose we want to download a webpage and print its HTML.

package main

import (
"fmt"
"io"
"net/http"
)

func main() {
url := "https://example.com"

resp, err := http.Get(url)
if err != nil {
    fmt.Println("Request failed:", err)
    return
}

defer resp.Body.Close()

if resp.StatusCode != http.StatusOK {
    fmt.Println("Unexpected status:", resp.Status)
    return
}

body, err := io.ReadAll(resp.Body)
if err != nil {
    fmt.Println("Failed to read response:", err)
    return
}

fmt.Println(string(body))
Enter fullscreen mode Exit fullscreen mode

}

The flow is now:

http.Get()
↓
http.Response
↓
resp.Body
↓
io.ReadAll()
↓
[]byte
↓
string

That's the foundation of our scraper.


  1. Extracting information from the response

Downloading the HTML is only half the job.

A scraper is useful because it extracts specific information.

For example, imagine the response contains:

Go Web Scraping

Learning Go is interesting.

We could perform very simple text processing with "strings".

html := `

Go Web Scraping

Learning Go is interesting.

`

if strings.Contains(html, "Go Web Scraping") {
fmt.Println("Title found")
}

However, this approach becomes fragile very quickly.

HTML is structured data, so for serious HTML parsing, a dedicated HTML parser is usually more appropriate.

For example, Go developers commonly use:

golang.org/x/net/html

This is not part of the Go standard library, but it works naturally with the standard library's HTTP and I/O packages.


  1. Parsing HTML with "golang.org/x/net/html"

After downloading a webpage, we can parse the HTML into a document tree.

Conceptually:

HTML Document
│
▼

│
┌────┴─────┐
▼ ▼

│
┌────┴────┐
▼ ▼

The parser allows us to navigate this structure instead of treating HTML as plain text.

A simplified example:

doc, err := html.Parse(resp.Body)
if err != nil {
fmt.Println(err)
return
}

Now "doc" represents the parsed HTML document.

From there, we can walk through the nodes and look for elements such as:





  1. Finding links

Links are particularly important when building a crawler.

Suppose we have:

About

Products

A crawler could find those "href" attributes and then visit them.

Conceptually:

Page A
│
├── /about
│
└── /products
│
▼
Page B

This is where web scraping starts becoming web crawling.

A scraper might extract information from one page.

A crawler can follow links and discover additional pages.


  1. "net/http/httptest" — testing the scraper

One of the most useful packages I encountered was:

net/http/httptest

Testing a scraper against a real website is not always a good idea.

The website could:

  • change its HTML
  • become unavailable
  • block requests
  • return different data
  • become slow

Instead, we can create a fake HTTP server.

Example:

server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
fmt.Fprintln(w, "

Hello

")
}))

defer server.Close()

Now we have a local test server.

We can send our scraper to:

server.URL

instead of a real website.


  1. Testing a scraper with a fake server

For example:

func TestScraper(t *testing.T) {
server := httptest.NewServer(http.HandlerFunc(
func(w http.ResponseWriter, r *http.Request) {
fmt.Fprintln(w, "

Hello

")
},
))
defer server.Close()

resp, err := http.Get(server.URL)
if err != nil {
    t.Fatal(err)
}

defer resp.Body.Close()

body, err := io.ReadAll(resp.Body)
if err != nil {
    t.Fatal(err)
}

if !strings.Contains(string(body), "Hello") {
    t.Fatal("expected Hello in response")
}
Enter fullscreen mode Exit fullscreen mode

}

Now our test doesn't depend on the internet.

We control exactly what the server returns.

That's extremely useful.


  1. Testing different HTTP responses

We can also test how our scraper handles errors.

For example:

server := httptest.NewServer(http.HandlerFunc(
func(w http.ResponseWriter, r *http.Request) {
http.Error(w, "Not Found", http.StatusNotFound)
},
))

Now our scraper receives:

404 Not Found

We can verify that our program handles that situation correctly.

We can also test:

200 OK
301 Redirect
403 Forbidden
404 Not Found
500 Internal Server Error

This is where testing becomes much more valuable than simply checking whether the scraper works once.


  1. Putting the packages together

At this point, the roles of the packages become clearer.

Package| Purpose
"net/http"| Send HTTP requests and receive responses
"io"| Read response data
"net/url"| Parse and construct URLs
"strings"| Process textual data
"net/http/httptest"| Test HTTP-based applications
"golang.org/x/net/html"| Parse HTML documents

They aren't competing packages.

They form a pipeline.

         URL
          │
          ▼
      net/url
          │
          ▼
      net/http
          │
          ▼
    HTTP Response
          │
          ▼
         io
          │
          ▼
      HTML data
          │
          ▼
    x/net/html
          │
          ▼
    Extract data
          │
          ▼
       strings
Enter fullscreen mode Exit fullscreen mode

  1. What I learned beyond the packages

The biggest lesson for me wasn't actually the individual packages.

It was understanding how they work together.

Before this, it was easy to think of web scraping as:

Download webpage
Extract information

But a more realistic system looks like:

URL
↓
HTTP request
↓
Timeout handling
↓
HTTP status validation
↓
Response body
↓
HTML parsing
↓
Data extraction
↓
Data validation
↓
Storage
↓
Tests

And once you start building a real scraper, additional concerns appear:

  • retries
  • timeouts
  • redirects
  • rate limiting
  • robots.txt
  • concurrent requests
  • duplicate URLs
  • malformed HTML
  • authentication
  • logging
  • error handling
  • data persistence

That's where a simple scraper starts turning into a real software engineering project.


  1. A small mental model

The way I'm currently thinking about Go web scraping is:

net/http
↓
"Talk to the website"

io
↓
"Read what the website sent"

net/url
↓
"Understand the address"

x/net/html
↓
"Understand the HTML structure"

strings
↓
"Clean and process text"

httptest
↓
"Test everything without depending on the real website"

Once you understand those responsibilities, the code becomes much easier to reason about.


  1. What's next?

My next step is to move beyond a basic scraper and explore how to build a more robust crawler in Go.

Some areas I want to explore are:

  1. HTTP clients and custom timeouts
  2. HTML tree traversal
  3. Extracting links and images
  4. URL normalization
  5. Concurrent scraping with goroutines
  6. Rate limiting
  7. Retry strategies
  8. Structured data models
  9. Testing edge cases
  10. Persisting scraped data
  11. Handling failed requests
  12. Building a complete CLI scraper

The interesting part is that the deeper I go, the more I realize that web scraping is not really about scraping.

It's about understanding how different parts of a software system communicate.

And Go's standard library provides many of the building blocks needed to understand and build those systems.

Top comments (0)