99.9% uptime since 1994. Zero cloud bills. Three servers running 50+ services with automated deployment, SSL termination, and zero-downtime updates.
That's my home network at i80.dk. After 30+ years of infrastructure work—migrating enterprise Kubernetes clusters, maintaining multi-datacenter services, coordinating 20+ person IT teams—I came home and built what I actually wanted: something simple, reliable, and free from vendor lock-in.
This isn't another "look at my homelab" post. This is the documented reasoning, hard lessons, and specific decisions behind a production network that costs me €1,165/year instead of €4,308 in cloud fees, while giving me more control and better understanding of every component.
The Problem I Was Solving
When you work in enterprise infrastructure, you get used to certain luxuries: managed Kubernetes, automatic SSL certificates, service meshes, observability stacks. Then you go home and realize AWS wants €450/month to replicate what three servers can do for the cost of electricity.
I had specific requirements:
- No cloud bills - Hardware I own, bandwidth I control
- Real automation - git push deploys services, nginx configs generated automatically
- Actual SSL everywhere - Not self-signed certificates and browser warnings
- Service discovery - Deploy a service, get a URL automatically
- Zero-downtime deploys - Update services without downtime
- Minimal complexity - No Kubernetes, no service mesh, no CNI plugins
The constraint: Do it with technology I could fix at 3 AM without documentation.
The Architecture That Works
The three-server design wasn't my first attempt. I started with a single beefy server running everything, then tried Kubernetes on three nodes, then Docker Swarm. Each iteration taught me something about what actually matters when you're the only operator.
The breakthrough came when I stopped asking "what would work at Google scale?" and started asking "what can I debug at 3 AM when I'm half asleep?" The answer surprised me: simpler tools with clear responsibilities.
Trinity (192.168.15.1) - Edge router and SSL termination
This is the only server that faces the internet, and that's deliberate. When you expose services to the public internet, you want a single point of entry you can secure, monitor, and understand completely. Trinity runs on an HP ProLiant MicroServer—not because it's powerful, but because it has multiple network interfaces. One connects to the internet, the other to my internal LAN via a dedicated network bridge.
Why dedicate hardware to this? Because mixing your network gateway with your application workloads is asking for trouble. If Nomad crashes during a deployment, I don't lose internet access. If I'm testing something risky on a VM, it can't take down my DHCP server. Separation of concerns isn't just good software design—it's good operational practice.
Trinity handles SSL termination because that's where the TLS handshake happens. I use Let's Encrypt wildcard certificates (*.i80.dk, *.zado.dk) which means every service gets proper HTTPS without individual certificate management. The nginx reverse proxy reads these certs from /infrastructure and routes traffic based on the domain in the request. One certificate, unlimited services.
The DHCP server runs here because it needs to be on the same Layer 2 broadcast domain as clients. You can't route DHCP requests through multiple network hops without special configuration. Trinity sits at the network edge with direct LAN access, making DHCP responses fast and reliable. When a new device boots, it gets an IP from Trinity within milliseconds.
Nomad (192.168.15.80) - Control plane
This server orchestrates everything. It's a Debian 12 VM running on a Dell R710 via Proxmox, and it hosts both Nomad (the orchestrator) and Consul (the service registry). I could have split these onto separate VMs, but that felt like over-engineering. They work together so closely that co-locating them simplified operations without sacrificing reliability.
Nomad's job is scheduling: when I deploy a service, Nomad decides where it runs and ensures it stays running. Consul's job is discovery: when a service starts, it registers in Consul, and other services can find it. The magic happens in consul-template, which watches Consul for changes and automatically generates nginx configs and PowerDNS records.
This server mounts Trinity's /infrastructure directory via NFS because that's where generated configs need to be written. When consul-template creates a new nginx config, it writes it to /infrastructure/consul-nginx/conf.d/, which Trinity reads directly. No rsync scripts, no ansible playbooks, no manual file copying. The filesystem is the API.
Autobox (192.168.15.124) - Worker node
This is where containers actually run. It's another Debian 12 VM on the same Dell R710, running Docker and the Nomad client agent. When Nomad schedules a job, it tells Autobox "run this container with these settings." Autobox pulls the image from my private registry, starts the container on a dynamically allocated port (30000-32000 range), and registers it with Consul.
Why dynamic ports? Because I don't want to manually track which port each service uses. Nomad picks an available port, Consul records it, and nginx routes to it automatically. If I need to scale a service to three instances, Nomad allocates three ports, Consul tracks all three, and nginx load-balances across them. Zero manual configuration.
The separation between Nomad (control plane) and Autobox (worker) mirrors production Kubernetes patterns, but it's simpler. Nomad doesn't run containers—it tells Autobox to run them. If Autobox crashes, Nomad detects it and marks all its services as unhealthy. If Nomad crashes, Autobox keeps running existing containers until Nomad comes back. This separation means I can restart either component without taking everything offline.
Both Nomad and Autobox run as VMs on Proxmox on the Dell R710 at stv.i80.dk, but I don't mention Proxmox in the architecture diagram because it's implementation detail. The network cares about Nomad and Autobox as logical components, not which hypervisor hosts them. This is the same reason I don't talk about which RAID level I use—it's infrastructure implementation, not network architecture.
The Deployment Flow That Changed Everything
I spent years in environments where deployments took hours. You'd commit code, create a pull request, wait for CI, get approval, merge to main, wait for another CI run, create a release, wait for someone to click "deploy" in a UI, wait for Kubernetes to pull the image and roll out pods, then manually update DNS entries and create JIRA tickets to update load balancer rules. By the time your code reached production, you'd forgotten what it did.
The breakthrough wasn't just making this faster—it was eliminating every manual step. When I git push to my Gitea instance, I don't touch anything else. No kubectl commands, no terraform apply, no "did someone update the DNS?" The code goes from my laptop to a production HTTPS URL in under a minute, fully automated.
This changes how you work. Instead of batching changes into big releases because deployments are expensive, you deploy every commit. Bug fix? Deploy it. New feature? Deploy it. Typo in the README? Who cares, it doesn't affect production, but if it did—deploy it. The friction between "I made a change" and "users see the change" is so low that deployment becomes a non-event.
Here's what actually happens when I git push:
1. Gitea Actions builds the image (30 seconds)
- name: Build Docker image
run: |
docker build -t registry.i80.dk/myapp:${{ github.sha }} .
docker tag registry.i80.dk/myapp:${{ github.sha }} registry.i80.dk/myapp:latest
2. Push to private registry (10 seconds)
docker push registry.i80.dk/myapp:${{ github.sha }}
docker push registry.i80.dk/myapp:latest
3. Deploy to Nomad (5 seconds)
nomad job run -var="image_tag=${{ github.sha }}" myapp.hcl
4. Nomad schedules to Autobox (2 seconds)
This is where orchestration happens. Nomad looks at the job specification and decides where to run it. Since I only have one worker node (Autobox), the decision is simple, but Nomad still checks: Does Autobox have enough CPU and memory? Are the required ports available? Is the container image accessible?
Once Nomad confirms Autobox can run the job, it allocates a dynamic port from the 30000-32000 range. This port allocation is automatic—I never specify ports in my job files. Nomad picks one, tells Autobox to start the container on that port, and records the allocation in Consul. If I scale the service to three instances, Nomad allocates three different ports and tracks all of them.
Autobox pulls the image from my private registry (registry.i80.dk), starts the container, and reports back to Nomad. The entire schedule-pull-start cycle takes about 2 seconds for small images, longer for multi-gigabyte containers. This is faster than Kubernetes because there's no CNI plugin to configure, no admission controllers to check, and no pod security policies to validate. Nomad just tells Docker "run this," and Docker does it.
5. Service registers with Consul (automatic)
The moment the container starts, Nomad tells Consul "this service is running at this address and port." This registration is automatic—I don't write registration code in my application. The Nomad job file includes a service block, and Nomad handles the Consul integration.
service {
name = "myapp"
port = "http"
tags = [
"traefik.enable=true",
"traefik.http.routers.myapp.rule=Host(`myapp.i80.dk`)",
"traefik.http.routers.myapp.tls=true"
]
check {
type = "http"
path = "/health"
interval = "10s"
timeout = "2s"
}
}
6. Consul-template generates nginx config (2 seconds)
Consul-template is the glue between service discovery and routing. It runs on the Nomad server, watching Consul for service changes. When "myapp" registers, consul-template sees it immediately and regenerates the nginx configuration file.
The template queries Consul's service catalog: "Give me all healthy instances of myapp." Consul responds with addresses and ports. Consul-template plugs this data into an nginx config template and writes the result to /infrastructure/consul-nginx/conf.d/services.conf.
This is why NFS matters. Consul-template runs on Nomad, but the config needs to land on Trinity. With the NFS mount, consul-template just writes to /infrastructure, and Trinity sees the change instantly. No SSH, no file copying, no synchronization delays.
upstream myapp_backend {
server 192.168.15.124:30450;
}
server {
listen 443 ssl http2;
server_name myapp.i80.dk;
ssl_certificate /infrastructure/wildcard.i80.dk.crt_fullchain.crt;
ssl_certificate_key /infrastructure/wildcard.i80.dk.key;
location / {
proxy_pass http://myapp_backend;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
}
access_log /var/log/nginx-consul/myapp.access.log;
error_log /var/log/nginx-consul/myapp.error.log;
}
7. Config watcher reloads nginx (1 second)
Trinity needs to know when consul-template writes a new config. I could configure consul-template to SSH into Trinity and run nginx -s reload, but that requires SSH keys, network access from Nomad to Trinity, and error handling if the SSH connection fails.
Instead, a simple bash script runs on Trinity as a systemd service, checking the config file's hash every second. When the hash changes, it runs systemctl reload nginx and saves the new hash. This approach is stateless, requires no network calls, and handles failures gracefully—if the reload fails, the script just tries again on the next iteration.
#!/bin/bash
FILE="/infrastructure/consul-nginx/conf.d/services.conf"
HASH_FILE="/tmp/services.conf.hash"
while true; do
CURRENT_HASH=$(md5sum "$FILE" | awk '{print $1}')
STORED_HASH=$(cat "$HASH_FILE" | awk '{print $1}')
if [ "$CURRENT_HASH" != "$STORED_HASH" ]; then
systemctl reload nginx
md5sum "$FILE" > "$HASH_FILE"
fi
sleep 1
done
8. Service accessible at https://myapp.i80.dk (total: 50 seconds)
The entire chain—from git commit to live HTTPS URL—completes in under a minute. No manual deployment steps. No waiting for approvals. No filing tickets to update load balancers. The code I wrote 50 seconds ago is now serving production traffic with automatic SSL termination, health checking, and service discovery.
This speed matters because it changes behavior. When deployment is expensive (slow, manual, error-prone), you batch changes. You wait until Friday afternoon to deploy a week's worth of work. You test manually because "I don't want to deploy this 10 times to fix bugs." You avoid refactoring because "it's not worth a deployment."
When deployment takes 50 seconds and is fully automated, you stop thinking about it. You write code, push it, and move on. Bug found? Fix and push. Feature idea? Try and push. The deployment pipeline becomes invisible, and you can focus on writing code instead of managing releases.
To put this in perspective: achieving this same flow with Kubernetes would require installing and maintaining 10+ components: the cluster itself, ArgoCD for GitOps, ingress controller, cert-manager, external-dns, and various operators. Each component adds complexity and debugging surface area. When something breaks at 3 AM, you're tracing through multiple controllers and CRDs to find the issue.
My stack? Six components total: Gitea Actions, Docker registry, Nomad, Consul, consul-template, and nginx. When something breaks, I check three places: Nomad logs, Consul state, or nginx logs. That's it.
Why NFS for Certificates Changes Everything
Most homelab setups handle configuration with either rsync scripts ("copy files to all servers every hour") or configuration management tools like Ansible ("run a playbook to update nginx configs"). Both approaches work, but they introduce latency and complexity. Changes don't take effect until the next sync runs, and you need to track which servers need which files.
I solved this with NFS—not because it's fancy, but because it's simple and solves three problems at once.
Trinity exports /infrastructure as a read-only NFS share (with one exception: Nomad can write to it). This directory contains certificates, generated nginx configs, and generated DNS records. Every server that needs this data mounts it over NFS, and changes are visible immediately.
Why this matters for consul-template:
Consul-template runs on the Nomad server, watching Consul for service changes. When a new service registers, consul-template generates an nginx config file. But here's the problem: Trinity runs nginx, not Nomad. How does the config get from Nomad to Trinity?
With NFS, consul-template just writes to /infrastructure/consul-nginx/conf.d/services.conf. Trinity has that directory mounted and symlinked into nginx's config directory. A systemd service watches for file changes and reloads nginx automatically. The entire flow—service registers → consul-template generates config → nginx reloads—takes 2-5 seconds, with zero manual intervention.
The alternative? I'd need to run consul-template on Trinity (more services to manage), or set up SSH keys so Nomad can scp files to Trinity (security complexity), or use Ansible to copy configs every minute (latency and cron jobs). NFS eliminates all of this.
Why this matters for certificates:
Let's Encrypt certificates expire every 90 days. Certbot on Trinity handles renewal automatically, but every other server needs the new certificates. How do you distribute them?
With NFS, there's nothing to distribute. Trinity stores certificates in /infrastructure/wildcard.i80.dk.crt_fullchain.crt, and every device that needs HTTPS mounts that directory. When certbot renews, the new certificate is instantly visible to all NFS clients. No distribution scripts, no ansible playbooks, no manual copying.
This enables something powerful: any device on my LAN can serve HTTPS with valid certificates. My Raspberry Pi running a dashboard? It mounts /infrastructure, points nginx at the wildcard cert, and serves https://dashboard.i80.dk with a real certificate. No self-signed warnings, no certificate management per-device. One wildcard cert, unlimited devices.
Why this matters for consistency:
The biggest win isn't speed or simplicity—it's that there's only one source of truth. When I check a certificate's expiration date, I look at /infrastructure on Trinity. That's the certificate every service uses. When I debug an nginx config issue, I look at /infrastructure/consul-nginx/conf.d/. That's the config Trinity is actually reading.
In systems with rsync or Ansible, you never know if a config file is up-to-date. Did the sync run? Did it fail? Is this server using the latest version? With NFS, if Trinity has it, everyone has it. There's no drift, no sync lag, no "let me check all the servers" debugging.
Certificate Management: Wildcard Certs for Everything
Why wildcard certificates matter:
Trinity runs certbot with Let's Encrypt to generate wildcard certificates (*.i80.dk, *.zado.dk, etc.). This isn't just for web services—these certificates enable:
1. Secure email (end-to-end encryption) - Mail servers use the same wildcard cert - SMTP/IMAP with proper TLS - No self-signed certificate warnings
2. Any service gets HTTPS automatically - Deploy new service → gets DNS record → nginx proxy uses wildcard cert - No per-service certificate generation - No certificate management per application
3. Internal services get proper HTTPS
- LAN devices mount /infrastructure via NFS
- Use wildcard cert for local HTTPS
- Browser shows proper TLS, no warnings
Certificate renewal is fully automated:
# On Trinity - runs automatically via systemd timer
certbot renew
# Certificates stored in /infrastructure/
/infrastructure/wildcard.i80.dk.crt_fullchain.crt
/infrastructure/wildcard.i80.dk.key
/infrastructure/wildcard.zado.dk.crt_fullchain.crt
/infrastructure/wildcard.zado.dk.key
Why this matters beyond web:
Most homelabs use self-signed certs or Let's Encrypt per-service. Wildcard certs + NFS export means: - Email servers get valid TLS - Internal dashboards work with HTTPS - APIs can enforce TLS without cert juggling - Mobile apps don't complain about certificates
The cost: €0. Let's Encrypt is free. The benefit: Professional-grade TLS everywhere.
Network Architecture
Here's how the pieces fit together:
Client: https://app.i80.dk"] Trinity["TRINITY
━━━━━━━━━━━
SSL Termination
Nginx reverse proxy
DHCP Server
PowerDNS
NFS Export /infrastructure"] Consul["CONSUL
━━━━━━━━━━━
Service Registry
Health Checking
KV Store
Consul-Template"] Nomad["NOMAD
━━━━━━━━━━━
Orchestration
Scheduling
Job Management"] Autobox["AUTOBOX
━━━━━━━━━━━
Nomad Client
Docker Engine
Containers"] Devices["DHCP DEVICES
━━━━━━━━━━━
Printers, IoT
Auto-registered"] LAN["LAN CLIENTS"] Internet -->|"HTTPS"| Trinity Trinity -->|"Proxy HTTP"| Autobox Trinity -.->|"Serves DHCP
Resolves DNS"| Devices Trinity -.->|"Serves DHCP
Resolves DNS"| LAN Nomad -->|"Deploy jobs"| Autobox Nomad -.->|"Queries service
state"| Consul Consul -.->|"Watches services
Generates nginx configs
Generates DNS records"| Trinity Devices -.->|"Registers
services"| Consul Autobox -.->|"Registers
services"| Consul Trinity -.->|"NFS mount
/infrastructure"| Consul Trinity -.->|"NFS certs"| LAN LAN -->|"Access via
Trinity or direct"| Trinity LAN -.->|"Direct access"| Autobox style Internet fill:#e1f5ff,stroke:#0288d1,stroke-width:3px style Trinity fill:#ffe0e0,stroke:#d32f2f,stroke-width:3px style Consul fill:#fce4ec,stroke:#c2185b,stroke-width:2px style Nomad fill:#e8f5e9,stroke:#388e3c,stroke-width:2px style Autobox fill:#fff3e0,stroke:#f57c00,stroke-width:2px style Devices fill:#f3e5f5,stroke:#7b1fa2,stroke-width:2px style LAN fill:#f1f8e9,stroke:#689f38,stroke-width:2px
Key insight: Trinity is the only server that faces the internet. Everything else is internal. DHCP triggers registration, Consul tracks state, consul-template generates configs, watchers reload services. Zero manual intervention.
Flow explanation:
- Solid lines: Direct traffic/data flow
- Dashed lines: Infrastructure/configuration flow
- Trinity terminates SSL and proxies HTTP to backends
- Consul is the service registry - everything registers here, consul-template watches and generates configs
- Nomad orchestrates deployments, queries Consul for service state
- Autobox runs containers, services auto-register in Consul
- DHCP devices get IP from Trinity, register in Consul, trigger config regeneration
- LAN clients can access via Trinity (HTTPS) or direct to services (HTTP)
Key insight: Consul is the source of truth. Nomad reads from it to make deployment decisions. Devices/services write to it. Consul-template (part of Consul) watches it and generates nginx/DNS configs.
Service Discovery Without Service Mesh
Consul provides service discovery without the complexity of Istio or Linkerd. When a service starts:
Service registers automatically (via Nomad integration)
{
"ID": "myapp-abc123",
"Name": "myapp",
"Address": "192.168.15.124",
"Port": 30450,
"Tags": ["traefik.enable=true", "HOST=myapp.i80.dk"],
"Check": {
"HTTP": "http://192.168.15.124:30450/health",
"Interval": "10s"
}
}
Consul-template queries Consul (every 2-10 seconds)
{{- range service "myapp" }}
upstream myapp_backend {
server {{ .Address }}:{{ .Port }};
}
{{- end }}
Nginx config generated (only healthy instances)
If health check fails, Consul marks service unhealthy, consul-template removes it from config, nginx stops routing. No manual intervention.
The DNS Integration Nobody Talks About
DHCP + Consul + DNS = fully automated network registration.
Why Trinity Runs DHCP (Not Nomad/Consul):
This isn't arbitrary—it's about network topology:
-
Broadcast Domain - DHCP operates at Layer 2/3 and requires being on the same broadcast domain as clients. Trinity sits at the network edge with a dedicated Gigabit NIC for LAN uplink.
-
Separation of Concerns - Trinity handles network infrastructure (routing, NAT, DNS, DHCP). Nomad/Consul handle application orchestration. If Trinity fails, you lose network anyway. If Nomad fails, network services continue.
-
Performance - DHCP requires fast responses to broadcasts. Trinity being directly on LAN provides minimal latency. Routing through additional hops adds delay.
-
Integration Pattern - Trinity runs DHCP and pushes device metadata to Consul. This "push" model is cleaner than orchestration trying to manage network services.
The Automation Chain:
When a new VM boots from the preseeded ISO or any device joins the network:
1. VM boots → requests IP via DHCP - Preseeded ISO contains hostname - DHCP server on Trinity receives request
2. DHCP assigns IP → triggers registration script
# /etc/dhcp/dhcpd.conf
on commit {
execute("/etc/dhcp/register-to-consul.sh", "i80.dk", hostname, ip_address, "commit", port);
}
3. Script registers device in Consul
curl -X PUT "http://192.168.15.80:8500/v1/agent/service/register" -d "{
\"ID\": \"dhcp-newvm\",
\"Name\": \"newvm\",
\"Tags\": [\"dhcp\", \"traefik.enable=true\"],
\"Address\": \"192.168.15.150\",
\"Port\": 80,
\"Meta\": {\"fqdn\": \"newvm.i80.dk\", \"ip\": \"192.168.15.150\"}
}"
4. Consul-template detects new service → generates DNS record
# /infrastructure/consul/trinity_powerdns_records.txt
i80.dk newvm A 192.168.15.150
5. DNS watcher sees file change → updates PowerDNS
pdnsutil replace-rrset "i80.dk" "newvm" "A" "192.168.15.150"
pdns_control reload
6. Consul-template generates nginx config (if service port is set)
upstream newvm_backend {
server 192.168.15.150:8080;
}
server {
listen 443 ssl http2;
server_name newvm.i80.dk;
ssl_certificate /infrastructure/wildcard.i80.dk.crt_fullchain.crt;
ssl_certificate_key /infrastructure/wildcard.i80.dk.key;
location / {
proxy_pass http://newvm_backend;
}
}
7. Config watcher reloads nginx
Result: VM boots → gets IP → registers in Consul → gets DNS record → gets nginx proxy (if needed) → accessible at https://newvm.i80.dk. Total time: ~10 seconds.
For services deployed via Nomad, same flow—except registration happens when container starts instead of VM boot.
Why Consul Instead of etcd or DNS
I evaluated three options for service discovery:
etcd + confd: More complex, requires separate consensus cluster, no health checking
CoreDNS + custom scripts: No health checks, manual service registration
Consul: Service discovery, health checking, KV store, single binary
Consul won because:
- Built-in health checking - Services that fail checks removed automatically
- Nomad integration - Service registration automatic on container start
- HTTP API - Easy to query from scripts (DHCP registration)
- DNS interface - Can query via DNS if needed (dig @localhost -p 8600 myapp.service.consul)
- Single binary - One systemd unit, no Raft cluster to maintain separately
The Nomad-Consul integration specifically removes manual work:
job "myapp" {
group "myapp-group" {
task "myapp" {
driver = "docker"
service {
name = "myapp"
port = "http"
# Service auto-registered in Consul when task starts
# Deregistered when task stops
# Health checks run automatically
}
}
}
}
Why Nomad Instead of Kubernetes
This decision surprised people. I've migrated 50+ applications to Kubernetes professionally, coordinated multi-datacenter K8s deployments, and spent years mastering its complexity. So why didn't I use it at home?
Because Kubernetes is built for a different problem. When you have multiple teams deploying hundreds of services, you need strong boundaries, fine-grained RBAC, and network isolation. You need admission controllers to enforce policies, pod security standards to prevent privilege escalation, and network policies to control traffic between namespaces. These features exist because Google, where Kubernetes originated, has thousands of engineers who might accidentally (or maliciously) break things.
But I'm the only engineer. I don't need protection from my teammates—I need something I can debug when things break.
Kubernetes has an enormous surface area. To run a simple web service, you need to understand: CNI (network plugins), CSI (storage drivers), ingress controllers, service meshes, admission controllers, RBAC policies, pod security standards, cert-manager for SSL, external-dns for DNS automation, and probably a dozen other components I'm forgetting. Each of these solves a real problem at scale, but they're all extra moving parts that can fail.
When something breaks in Kubernetes at 3 AM, you're troubleshooting a distributed system with 15+ components. Is it the ingress controller? Did cert-manager fail to renew? Is the CNI plugin misconfigured? Did a network policy block the traffic? Is the pod security policy rejecting the deployment? You have to know all of these systems deeply to diagnose problems quickly.
Nomad, by contrast, has four concepts: jobs (what to run), groups (how to co-locate), tasks (the actual container), and drivers (how to run it). That's it. No CNI—Nomad uses the host network. No CSI—Nomad uses host volumes or NFS. No ingress controller—I use nginx directly. No cert-manager—I use certbot. When something breaks, I have four places to look instead of fifteen.
Example Kubernetes deployment:
# deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: myapp
spec:
replicas: 1
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
spec:
containers:
- name: myapp
image: myapp:latest
ports:
- containerPort: 8080
---
# service.yaml
apiVersion: v1
kind: Service
metadata:
name: myapp
spec:
selector:
app: myapp
ports:
- port: 80
targetPort: 8080
---
# ingress.yaml
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: myapp
annotations:
cert-manager.io/cluster-issuer: letsencrypt
spec:
tls:
- hosts:
- myapp.i80.dk
secretName: myapp-tls
rules:
- host: myapp.i80.dk
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: myapp
port:
number: 80
Equivalent Nomad deployment:
job "myapp" {
datacenters = ["dc1"]
group "myapp-group" {
count = 1
network {
port "http" {
to = 8080
}
}
task "myapp" {
driver = "docker"
config {
image = "registry.i80.dk/myapp:latest"
ports = ["http"]
}
service {
name = "myapp"
port = "http"
tags = [
"traefik.enable=true",
"traefik.http.routers.myapp.rule=Host(`myapp.i80.dk`)",
"traefik.http.routers.myapp.tls=true"
]
}
}
}
}
Nomad: 1 file, 30 lines, no additional components.
Kubernetes: 3 resources, 50+ lines, requires cert-manager and ingress controller.
When you're the only operator and something breaks at 2 AM, Nomad's simplicity matters.
The Cost Reality: Cloud vs. Self-Hosted
My infrastructure runs 50+ services. Here's the monthly cost comparison:
Cloud (AWS equivalent):
- 3x t3.medium instances: €108/month
- 500GB EBS storage: €50/month
- Load balancer: €16/month
- Data transfer (2TB/month): €180/month
- Route53 + ACM: €5/month
- Total: €359/month (€4,308/year)
Self-hosted (actual costs):
- Electricity (3 servers, 24/7): €45/month
- Internet (1Gbit/s fiber): €50/month (company-paid, but included for comparison)
- Domain registrations (4-5 domains): ~€25/year
- Total: €95/month (€1,165/year) (€540/year if internet excluded)
Savings: €3,143/year (€3,768/year if internet excluded)
But the real cost difference shows over 5 years:
Cloud (5 years):
- Services: €21,540
- No hardware ownership
- Vendor lock-in increases costs over time
- Total: €21,540+
Self-hosted (5 years):
- Services: €5,825
- Hardware: €2,500 (servers purchased upfront)
- Upgrades/replacements: €1,000
- Total: €9,325
5-year savings: €12,215
The numbers don't include:
- Cloud cost increases (3-5% annually)
- Egress fees when you switch providers
- Premium support when things break
- Development time debugging vendor-specific issues
The Three-Server Minimum (Two Physical Boxes)
Why three servers? Can't you do this with one?
One server: Single point of failure. Update = downtime. Hardware failure = complete outage.
Two servers: Split-brain potential. Can't do quorum consensus. Maintenance still requires downtime.
Three servers: Quorum possible. Maintenance without downtime. Failover capability.
My actual setup:
- Trinity - Separate HP box (network gateway must be separate)
- Dell R710 - Runs both Nomad and Autobox as separate VMs/containers
In practice:
- Trinity down: LAN services continue, external access lost
- Nomad down: Can't deploy new services, existing services run
- Autobox down: Services stop, can deploy to different node
- Dell R710 down: Both Nomad and Autobox lost, but Trinity keeps network running
The separation between Trinity (network infrastructure) and the Dell R710 (application workloads) is the critical split. Within the Dell, having Nomad and Autobox separate provides orchestration flexibility.
What I'd Do Differently
After running this for 2+ years, here's what I learned the hard way:
1. Log aggregation from day one
I have 50+ services logging to individual files on Autobox. When debugging issues, I SSH and grep logs. This works but doesn't scale. Should have set up Loki + Grafana on day one.
Better approach:
task "myapp" {
driver = "docker"
config {
logging {
type = "syslog"
config {
syslog-address = "udp://192.168.15.80:514"
tag = "myapp"
}
}
}
}
Central logging beats SSH + grep.
2. Metrics and monitoring
Currently: Manual checks via Consul UI and Nomad UI
Should be: Prometheus + Grafana with alerts
Missing metrics: - Container CPU/memory over time - Nginx request rates per service - Disk space trends (would have caught the DHCP log issue) - Certificate expiration warnings
3. Automated backups
I backup configs manually. Should be:
#!/bin/bash
# Daily backup script
BACKUP_DIR="/backup/$(date +%Y%m%d)"
# Trinity configs
tar -czf "$BACKUP_DIR/trinity.tar.gz" \
/etc/nginx/ \
/etc/dhcp/ \
/infrastructure/
# Nomad configs
tar -czf "$BACKUP_DIR/nomad.tar.gz" \
/etc/consul-template/ \
/etc/nomad.d/
# Service data (SQLite DBs, etc.)
rsync -av /opt/data/ "$BACKUP_DIR/data/"
# Push to remote
rsync -av "$BACKUP_DIR/" backup-server:/backups/i80/
4. Testing procedures before emergencies
I documented disaster recovery but never tested it. When Trinity's disk filled, I realized: - No monitoring caught it - No alerts warned me - Recovery documentation untested
Should have: - Quarterly DR tests - Automated health checks with alerts - Documented runbooks tested under stress
The Infrastructure-as-Code Reality
Everything should be in git. Here's what actually is:
In git:
- Application code (Gitea)
- Nomad job files (in app repos)
- Consul-template templates (versioned on Nomad)
- Gitea Actions workflows (
.gitea/workflows/) - Preseeded Debian ISO (auto-registers via DHCP)
Automated but not in git:
- Proxmox VMs (stv.i80.dk) - Automated deployment via preseeded ISO
- DHCP registration (devices auto-register on network join)
Manual (one-time setup):
- Trinity base install (Debian, nginx, PowerDNS, DHCP server)
- Initial systemd unit files (config watchers, DHCP hooks)
The reality: About 99% is fully automated. The only manual work left is Trinity's initial base install—a one-time setup I did years ago.
What works well:
- New VMs: Boot from preseeded ISO, automatically get IP, DNS, register in Consul, and update DHCP config on first boot
- New services:
git push→ builds → deploys → gets HTTPS URL - Certificates: Auto-renew and distribute via NFS
- DHCP reservations: Preseeded ISO updates
dhcpd.confon first boot
What was manual (now a distant memory):
- Trinity initial setup—done once in 1994, hasn't needed touching since the current architecture in 2022
- Everything else? Automated.
For Someone Starting Today
You're considering building something similar. Here's the minimum viable approach:
Start simple:
- 1 server running Nomad + Consul (combined mode)
- Docker as the only driver
- Manual nginx config for first few services
- Learn the components before automating
Scale when it hurts:
- nginx configs painful? Add consul-template
- Manual deploys annoying? Add git-based CI/CD
- One server risky? Add second node
- Need external access? Add edge proxy (separate box for network gateway)
Don't overbuild:
- Skip Vault until you have actual secrets management needs
- Skip service mesh until service-to-service auth matters
- Skip log aggregation until grep feels painful
Hardware recommendations:
Option 1 - Budget start (~€500):
- 1x Intel NUC or similar (Nomad + Consul)
- Use your existing router for network (no Trinity equivalent needed yet)
Option 2 - My setup (~€800):
- 1x HP ProLiant MicroServer or similar with multiple NICs (network gateway)
- 1x Dell PowerEdge R710 or similar server (run Nomad + Autobox as VMs)
- Benefits: Proper network separation, production-grade reliability
Option 3 - Enterprise-lite (~€2000):
- 2x dedicated servers for Nomad cluster
- 1x dedicated network gateway
- 1x NAS for shared storage
- Benefits: True HA, no single points of failure
The Bottom Line
This setup gives me:
- 50+ services running on 3 servers
- 99.9% uptime over 2+ years
- Sub-minute deployments from git push to production
- Zero cloud bills (€3,128/year savings vs. AWS)
- Automatic SSL for all services
- Service discovery without manual config
- DNS integration for every device
- Complete control over every component
The trade-offs:
- I'm the only operator (no team to share on-call)
- Hardware failures are my problem
- No managed services safety net
- Learning curve for Nomad, Consul, nginx, PowerDNS
- Physical access required for bare metal issues
What's missing (and how to add it):
- DDoS protection - Not included. Add Cloudflare (free tier works) in front of your IP
- Load balancing - Single Trinity handles all traffic. Add HAProxy or second Trinity for HA
- Geographic distribution - Single location. Add Cloudflare for CDN or secondary sites
- Managed WAF - No application firewall. Cloudflare or AWS Shield if needed
- 24/7 monitoring - Manual checks only. Add UptimeRobot or similar (free tier available)
For production workloads facing the public internet, I'd recommend Cloudflare's free tier as minimum—gives you DDoS protection, CDN, and SSL without changing your setup. Point your domain to Cloudflare, Cloudflare points to your IP. Zero cost, significant protection.
But after 30+ years in infrastructure, I'll take the trade-off. When something breaks, I know exactly where to look. When I need a new feature, I know exactly what to change. When costs come up, they're electricity and internet—not surprise AWS bills.
The best infrastructure is the one you understand completely.
Technical Specifications
For reference, here are the exact versions and configs that run this setup:
Trinity:
- Debian testing (trixie/sid)
- nginx 1.26.3
- PowerDNS Authoritative 4.9.4
- PowerDNS Recursor 5.2.2
- ISC-DHCP Server 4.4.3-P1
Nomad Server:
- Debian 12.12 Bookworm
- Nomad 1.10.5
- Consul 1.21.5
- consul-template 0.41.2
Autobox:
- Debian 13.1 Trixie
- Docker 26.1.5
- Nomad 1.10.5 (client)
- Consul 1.21.5 (client)
Network:
- Internal: 192.168.15.0/24
- External: 87.52.104.155 (IPv4 only)
- Internet: 1Gbit/s fiber
Physical Hardware:
- Trinity: HP ProLiant MicroServer with multiple NICs
- stv.i80.dk: Dell PowerEdge R710 (runs Nomad + Autobox as VMs via Proxmox)
Storage:
- Trinity: 256GB SSD
- Dell R710: Shared storage pool (Nomad + Autobox VMs)
About the author: Henrik Jess is a DevOps engineer with 30+ years of infrastructure experience. He's maintained self-hosted infrastructure with 99.9% uptime since 1994 and has professionally migrated 50+ applications to Kubernetes, coordinating teams across multiple datacenters. He runs i80.dk, a completely self-hosted infrastructure serving 50+ services from three servers in his home office.
Need help with your infrastructure? Whether you're looking to escape cloud vendor lock-in, build a similar setup, or need guidance on self-hosted infrastructure, reach out. Available for consulting and architecture discussions. Contact: web.i80.dk/contact