So far every variable in these posts has held a single value. Real programs deal with collections: a list of customers, a record for each student. Python has four built-in containers for that: lists, tuples, dictionaries, and sets.
Lists and dictionaries especially are the backbone of what comes next. A pandas DataFrame can be built from a dictionary of lists, where each key is a column, and every API response or JSON record arrives in Python as a dictionary.
Lists: ordered and changeable
A list keeps items in order, and you can add, remove, or change them whenever you like.
fruits = ["mango", "banana", "pineapple"]
print(fruits[0]) # mango
print(fruits[-1]) # pineapple
print(fruits[0:2]) # ['mango', 'banana']
fruits.append("orange")
print(len(fruits)) # 4
Positions start at 0, and -1 is always the last item. Lists also come with methods for changing them:
| Method | What it does |
|---|---|
append(x) |
Adds x to the end |
insert(i, x) |
Inserts x at position i
|
extend(list2) |
Adds every item from another list |
remove(x) |
Removes the first x it finds |
pop() / pop(i)
|
Removes and returns the last item, or the one at i
|
index(x) / count(x)
|
Position of x / how many times it appears |
sort() / reverse()
|
Sorts or reverses the list in place |
register = ["Bob", "John", "Faith", "Mercy", "Wanjiku"]
register.insert(0, "Amina")
register.extend(["Brian", "David"])
print(register.pop()) # David
register.remove("John")
print(register.index("Faith")) # 2
marks = [78, 45, 92, 60]
print(sorted(marks)) # [45, 60, 78, 92] new list, marks unchanged
marks.sort(reverse=True)
print(marks) # [92, 78, 60, 45] marks itself changed
Most of these methods change the list in place and return None, so marks = marks.sort() leaves you with None. Use sorted(marks) when you want a new list back.
Tuples: a list that's locked
A tuple can't be changed after it's created, which is exactly right for data that should never move: GPS coordinates, RGB colour codes, a row from a database.
nairobi = (-1.2921, 36.8219)
latitude, longitude = nairobi # unpacking
print(latitude, longitude) # -1.2921 36.8219
nairobi[0] = 0 # TypeError: 'tuple' object does not support item assignment
Dictionaries: values with labels
With a list you find things by position. With a dictionary you find them by a label, called a key, which is far more readable when each value means something different. It's the natural fit for student records, phone books, configs, and JSON.
student = {"name": "Amina", "age": 17, "class": "Form 3"}
print(student["name"]) # Amina
student["city"] = "Nairobi" # add a key
student["age"] = 18 # update a key
for key, value in student.items():
print(f"{key}: {value}")
print(student.get("phone", "not provided")) # not provided
student["phone"] would crash with a KeyError since that key doesn't exist, so .get() with a fallback is the safer choice when you're not sure.
Sets: unique items only
A set drops duplicates automatically and doesn't care about order, which makes it good for removing repeats and comparing groups.
cities = ["Nairobi", "Mombasa", "Nairobi", "Kisumu", "Mombasa"]
print(sorted(set(cities))) # ['Kisumu', 'Mombasa', 'Nairobi']
morning = {"Amina", "Brian", "Cynthia"}
evening = {"Brian", "David"}
print(morning & evening) # {'Brian'} in both
print(sorted(morning | evening)) # ['Amina', 'Brian', 'Cynthia', 'David'] in either
Which one should I reach for?
| Container | Ordered? | Changeable? | Duplicates? | Best for |
|---|---|---|---|---|
| List | Yes | Yes | Allowed | A sequence you add to or loop over |
| Tuple | Yes | No | Allowed | Fixed data like coordinates |
| Dictionary | Yes | Yes | Unique keys | Looking things up by label |
| Set | No | Yes | Not allowed | Removing repeats, checking membership |
If position matters, use a list. If a label matters, use a dictionary. If it must never change, use a tuple. If you only care about uniqueness, use a set.
Nesting them: a small grade book
The real power shows up when you combine them. Here a dictionary maps each student to a list of marks:
grades = {
"Amina": [78, 85, 90],
"Brian": [92, 55, 71],
"Cynthia": [49, 59, 71],
}
for student, marks in grades.items():
average = sum(marks) / len(marks)
print(f"{student}: {average:.1f}") # Amina: 84.3, Brian: 72.7, Cynthia: 59.7
Where this showed up in my project
The data for my Bundle Purchase Simulator is this kind of nesting. The catalogue is a dictionary where each category maps to a list, and each bundle in that list is its own dictionary:
bundles = {
"Data": [
{"name": "50MB - Daily", "price": 5, "validity": "24 hours"},
{"name": "500MB - Weekly", "price": 50, "validity": "7 days"},
{"name": "2GB - Monthly", "price": 200, "validity": "30 days"},
],
# "SMS" and "Minutes" follow the same pattern
}
print(bundles["Data"][1]["name"]) # 500MB - Weekly
print(bundles["Data"][1]["price"]) # 50
Each choice fits the job. Categories are looked up by name, so a dictionary. Bundles within a category are shown in numbered order, so a list. Each bundle has a name, a price and a validity that mean different things, so a dictionary again. Read left to right, bundles["Data"][1]["price"] is the Data category, its second bundle, then the price.
What I took from it
Nothing here is hard on its own. The skill is choosing well. Once I stopped squeezing everything into lists and gave each record labelled fields, the code started explaining itself.
Top comments (1)
Some comments have been hidden by the post's author - find out more