Python for Data Engineers

Installing Python packages with pip

Constantin LunguUpdated 1 min read

Photo by Evan Krause on Unsplash

Here are the basics you need to know about installing Python packages.

Apart from the built-in modules that come by default with the Python installation (the Standard Library), you need to install third-packages packages to be able to use them in your code (and of course import them).

You do so with a package manager. Perhaps the most widely known is pip, but there are other options like poetry or uv.

You can check more information about Python packages at pypi.org

Here's a quick list of common use cases:

➡ pip install pandas # installs the package and its dependencies

➡ pip uninstall pandas # removes the package

➡ pip install --upgrade pandas # upgrades a package

➡ pip list # lists installed packages

➡ pip freeze > requirements.txt # saves the list of installed packages to a file so you can recreate the environment with the same packages next time you need it

program.py:

import pandas as pd

data = [{'a': 1, 'b': 2}, {'a': 3, 'b': 4}]

df = pd.DataFrame(data)

df.head()

requirements.txt:

numpy==2.0.0
pandas==2.2.2
python-dateutil==2.9.0.post0
pytz==2024.1
six==1.16.0
tzdata==2024.1

Terminal output: source .venv/bin/activate adds a (test) prefix to the prompt and deactivate removes it; the user and host name are hidden.


Enjoyed this? Here are some related articles you might find useful: