neuralgia.

lessons from the trenches.

Python boilerplate.

Listing down some of the design decisions I am making for creating my own Python boilerplate. The use case is to create a repository that I can easily clone and hit the ground running by coding application logic instead of worrying about setting up CI/CD pipelines, pre-commit hooks, writing Dockerfiles, etc.

Python

I never had to worry about Python versions, since I do not usually create Python packages for distribution. Most of my work revolves around putting machine learning research code into production, creating specialized tools for personal use, etc. For this reason, I would usually pick a Python version and test functionality only for that specific version.

However, noting the differences between Python versions with regards to type-hinting multiple types amongst other improvements since Python 3.8, it would make sense to update as much as possible. Python 3.12 seemed stable enough at the time of writing without breaking too much code from earlier Python versions while having some major improvements.

Cookiecutter

I learned about `cookiecutter` while doing my apprenticeship and thought it was a good tool especially for the purpose of creating boilerplate. I used to provide some level of customizing project names via the use of configuration files in boilerplate code, but realized that most of the time it is contingent on people taking the time to modify the configuration files themselves.

Furthermore, it did not make sense to provide static project-level variables via configuration files, as they are fixed at project creation and not meant to be changed, thus the need for a template engine.

Cookiecutter provides the following solutions:

  1. It is possible to run cookiecutter on a repository remotely, removing the need for cloning a repository and renaming project folders.
  2. Cookiecutter removes the need for configuration files for project-level variables via user input prompts. This ensures nobody escapes the menial task of ensuring all project-level variables are set.
  3. Because it is a template engine, we can use it on files that are not necessarily Python scripts. This includes Dockerfiles and YAML configuration files.

Endpoints

The name “endpoints” is a little fuzzy, but essentially there are 3 types I decided to include in my boilerplate:

  1. API: Just a generic FastAPI template (enhanced via the use of cookiecutter to fill project-level variables.
  2. CLI: Usual Python script execution.
  3. Agent: Agent-based workflow, monitoring folders for changes, etc.

The idea is to eventually include Python package as one of the available endpoints. However that is new territory for me and I have no use for it at this point.

There is also something I learned while implementing `budge`; the idea is to create a single handler function that will kickstart all application-related logic (similar to a main() function used in most Python scripts). All endpoints will point to this function.

This consolidates the application logic flow into a single point, reducing the number of tests I have to run while still keeping the application logic generalized and ready to be integrated into other endpoints.

Testing & CI/CD

I decided to use Nox as a intermediate layer between GitLab/GitHub and my own software-level checks, formatting and tests. Using Nox to consolidate whatever I need to run for my unit/integration tests and my CI/CD pipeline would simplify the CI/CD pipeline definitions I need to create for GitLab and GitHub, with the added advantages:

  1. Nox is configured and executed via Python code, instead of YAML configurations, bash, etc. This reduces the amount of non-Python scripts and files and allows for my testing and CI/CD processes to be tested as well.
  2. It reduces the complexity of running my test cases, especially against multiple Python versions, different configurations, different environments, etc. This also includes computational complexity as Nox caches the Python environment used to run everything that is needed to be ran.
  3. It is perfectly suited for people who produce Python packages, who need to test their code against multiple Python versions, or have special configurations that may be difficult to implement as part of the CI/CD pipeline.

Leave a Reply

Your email address will not be published. Required fields are marked *