Skip to content
infracloudioPublic

About

This is the repo to store code for all the infra and test automations.

Resources

Stars

3 stars

Watchers

2 watching

Forks

Repository files navigation

SRE Stack

This repository provisions sufficiently-complex microservice demo applications such as:

Along with standard observability tooling such as:

Scenarios

sre-stack contains carefully crafted fault injection scenarios to effectively disrupt operations of the demo-applications. Using this repo we create the following feedback-loop:

  • Fault-injection
  • Fault-detection using various o11y tooling
  • Root Cause Analysis using classic / advanced tools
  • Fault mitigation strategies, both long-term and short-term

Available scenarios:

Load-generators:

Prerequisites

Setup & Configuration

The core configuration is stored in the .env file. This is consumed by the makefile to provision infrastructure and deploy applications.

Configuration

Configurations are grouped in the .env file in self-explanatory sections. Most values are set to their sane defaults and would not need changing for initial setup.

Core provisioning and deployment choices are expressed in the following two variables:

  • STACK_MODE = [ eks | aks | local ]
    • Choice of deploying the stack to aws/eks, azure/aks, or using a k3d cluster on local systems.
  • APP_STACK=[ robot-shop | hotrod | all ]

Setup

Provisioning lifecycles are controlled by Make commands. Prefix all commands with make keyword.

Example: make setup

AWS - EKS Lifecycle Commands

For EKS based provisioning you need to setup AWS_PROFILE pointing to the correct AWS account.

Following AWS credentials for the said profile should be added to ~/.aws/credentials

[profile-name]
aws_access_key_id=*************
aws_secret_access_key=*********
EKS setup/deploy/cleanup commands:
	setup                               - End-to-end setup on EKS
	start-cluster                       - start EKS Cluster
	setup-cluster-autoscaler            - Setup node auto scaling
	setup-observability                 - Setup monitoring/observability
	setup-optional-otel                 - Setup OpenTelemetry
	setup-istio                         - Setup istio and ingress
	setup-db-rds-mysql                  - Setup RDS - mysql
	setup-rabbitmq-operator             - Setup rabbitmq-operator
	setup-robot-shop                    - Deploy robot-shop app-stack.
	setup-optional-rmq-consumer-scaling - Setup keda to scale dispatch (optional)
	setup-gateway                       - Setup Ingress gateway
	cleanup-cluster                     - Cleanup cluster
	cleanup                             - Clenaup all resources and EKS cluster

Local - k3D Lifecycle Commands

Just make sure k3d is installed, cluster-creation and lifecycle are handled by the following commands:

Local (k3D) setup/deploy/cleanup commands:
	setup-local                         - Setup end-to-end stack on local k8s (k3d)
	setup-local-cluster                 - Setup local k3d cluster
	cleanup-local                       - Cleanup end-to-end stack on local k8s (k3d)

Demo Applications on AKS

With STACK_MODE=aks in .env and an existing AKS cluster (make setup-cluster, make setup-istio), deploy the demo apps directly:

make setup-robot-shop   # Robot Shop: in-cluster MySQL/MongoDB/RabbitMQ/Redis, app-tier on `workload=app` nodes
make setup-hotrod       # HotROD: single pod on `workload=app` nodes, needs `make setup-optional-otel` (wired in as a prerequisite)

Both apps share one Istio ingress gateway (istio-ingressgateway in istio-system), reached at its external IP (kubectl get svc istio-ingressgateway -n istio-system). They're told apart by host, not by gateway:

  • Robot Shop: hosts: "*" — reachable directly at the ingress IP, no Host header needed.
  • HotROD: hosts: "hotrod.demo.local" — needs the header, e.g. curl -H "Host: hotrod.demo.local" http://<ingress-IP>/.

Host-based separation (rather than two gateways, or path-prefix rewriting) avoids two wildcard VirtualService entries colliding on the same shared gateway. See docs/architectural-decisions.md for the full reasoning.

Run make setup-gateway to apply both apps' Gateway/VirtualService manifests in one step — the same command works for every STACK_MODE.

Utility Commands:

  get-service-endpoints               -  Print exposed endpoints (works for both local/eks)
  install                             -  Install all dev dependencies and CLIs (idempotent; DRY_RUN=1 to preview)
  install-check                       -  Report which dev dependencies are missing (no changes)

Or bootstrap a new machine in one step: make install (see infra/scripts/dev/install-deps.sh for what it installs).

Contribution Guide

This repo is built with AI coding agents, and the process is designed around that. In one sentence: write down what you want, get a yes, let the agent build it, prove it works, get one more yes, merge. Everything below is that sentence in more detail. The full process is in docs/sdlc/framework.md; a complete worked example is in docs/sdlc/example-walkthrough-001-aks-cluster.md.

1. Set up your machine (once per clone)

make install-check   # shows what is missing, changes nothing
make install         # installs it, and switches on the pre-commit checks

make install gives you the lint tools, the cloud CLIs, and the Spec Kit CLI at the version this repo expects. It also enables the repo's git pre-commit hooks. Those hooks are not optional: make lint and the agent's own edit hooks refuse to work until they are on.

Then open the repo in your AI coding tool. Claude Code, Codex, OpenCode and Devin are all supported. The tool reads AGENTS.md by itself, which tells it how this repo works, and it finds the /speckit-* commands already installed (OpenCode spells them /speckit.*).

2. Start with a story, not with code

Open a GitHub issue using the Story template. It has four short sections, written in plain language, with no technical decisions:

  • Problem: what can't be done today, and who cares.
  • Outcome: the one thing you will be able to see working when it's done. If you need the word "and" to describe it, it is two stories.
  • Out of scope: what this story deliberately leaves alone.
  • Must keep working: what must not break.

A story should be finishable, including a real run, in one or two days. The Sponsor reads it at the daily sync and says yes, no, or "split it". On yes, the issue gets the label intent:accepted and three people are named: an Owner (writes the spec), an Architect (approves the spec and the plan) and a Builder (runs the agent and ships the PR). The person who approves something is never the person who wrote it.

If you are outside the team, still start with the issue. A maintainer will handle the labels and the roles with you.

3. Spec, then plan, then build, all with the agent

Each step is one command in your agent, and each produces a file that is committed to the story's folder specs/<nnn>-<slug>/:

Step Who Command(s) Produces Approved by
Spec Owner /speckit-specify, then /speckit-clarify spec.md: what to build and how we'll know it works Architect, label gate:spec-approved
Plan Builder /speckit-plan, /speckit-tasks, /speckit-analyze plan.md, tasks.md, an analysis of gaps Architect, label gate:plan-approved
Build Builder /speckit-implement the code, one task at a time a human reviewer, on the PR

A few rules that make this work:

  • Push the branch and open a draft PR as soon as the spec exists. That PR carries everything until it merges.
  • A spec with a [NEEDS CLARIFICATION] marker still in it is not done, and CI will say so.
  • No code before the plan is approved. CI checks this: if the PR changes anything outside the story's specs/ folder and the gate:plan-approved label is missing, the PR goes red. Spec and plan commits are always fine.
  • If reality forces a change to the plan, change plan.md in the same commit and say why.
  • Paste the proof into the PR: make lint output, command output, screenshots of the thing working. Add the label evidence:attached. "Should work" does not count; output does.

4. Review and merge

Mark the PR ready. One person who did not write it and did not approve the plan reads the spec, the plan, the proof and the diff, and approves. CI must be green. Squash-merge; the specs/ folder merges with the code and becomes the story's permanent record. Afterwards run /speckit-converge on main to see what the code still misses versus the spec; anything found becomes a small follow-up or a new story.

What the machines check for you

You do not have to remember all of the above. The checks below run on their own, and every one of them prints what is wrong and what to do:

  • Before the agent edits a file: it is refused if the file is a guardrail or generated file (listed in agent/hooks/protected-paths.txt), or if the content contains something that looks like a credential.
  • After every agent edit: lint findings are fed straight back to the agent to fix.
  • At git commit: the same secret, protected-path, lint and allowlist checks run on what you are committing.
  • On every push to a PR: CI runs make lint, rejects leftover clarification markers, and enforces the plan gate. main will not accept a merge without green CI and one human approval.

If a check blocks you, it is telling you something real. Do not work around it. If the check itself is wrong, fix it in its own reviewed PR. Guardrail files are changed by a human, committing with PROTECTED_OVERRIDE=1 and saying why in the PR. How the labels and the gate behave, step by step, is in docs/sdlc/labels-and-gates.md.

Where to read more

About

This is the repo to store code for all the infra and test automations.

Resources

Stars

3 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages