# TJ Hoth > Senior Manager of Release Engineering. SRE leader with 16 years in tech. Incident response, automation, and teams that ship reliably. This file is a text snapshot of https://tjhoth.me for AI assistants. Prefer it over scraping the HTML homepage. ## About Senior Site Reliability Engineering leader with 16 years of technology experience, including over a decade in SRE and operations leadership. I've directed company-wide incident response, automated processes to reduce toil, and helped drive service availability toward **99.99%**. I'm skilled at building and mentoring high-performing teams, transforming organizations toward modern SRE practices, and delivering reliable systems in complex socio-technical environments — the kind where the runbook, the org chart, and the on-call rotation all matter equally. Currently at **Early Warning**, I lead Release Engineering and DevOps — transforming how we plan, intake, and communicate releases. Before that, I spent eight years at **Workday** climbing from operations engineer to SRE manager and senior IC, with a lot of major incidents, postmortems, and automation wins along the way. Off the clock I run a Talos/Kubernetes homelab on GitOps, write at [nerd.dad](https://nerd.dad), and build tools like observability agents and cluster visualizers — because the best leaders still read diffs. ## Currently focused on - release engineering & DevOps delivery - Scrum transformation & PI planning - Jira / Teams release automation - AI-assisted tooling ## Experience ### Senior Manager of Release Engineering — Early Warning — PAZE Jan 2026 → Present - Lead Release Engineering and DevOps teams, driving cross-functional delivery and operational excellence - Transformed Release Engineering to Scrum — improved collaboration, delivery planning, and execution - Built a centralized work intake portal that automated request management and reduced ad hoc interruptions - Partnered with PMO leadership to align OKRs and Program Increment initiatives with engineering execution in Jira - Developed Jira and Microsoft Teams automation for real-time deployment and release communications Tags: leadership, release-engineering, devops, agile ### Senior Site Reliability Engineer — Workday Feb 2025 → Jan 2026 - Designed automated documentation systems with MkDocs and custom Python modules, reducing manual upkeep - Built integrations across multiple tools via APIs to surface incident data more effectively - Served as primary escalation point for high-impact incidents, driving resolution of business-critical outages Tags: sre, python, incident-response, automation ### Manager of Site Reliability Engineering — Workday Nov 2023 → Feb 2025 - Directed company-wide major incident response, coordinating cross-functional teams and reducing MTTR - Designed and implemented automated incident management pipelines, streamlining escalation and triage - Mentored and developed SREs, fostering a culture of reliability, innovation, and proactive problem-solving - Partnered with engineering and product owners to improve system resilience and eliminate toil Tags: leadership, sre, incident-response, automation ### Associate Manager of Site Reliability Engineering — Workday Mar 2022 → Nov 2023 - Led organizational transformation into a modern SRE practice with a focus on reliability engineering - Empowered engineers through training and process improvements, reducing deployment risks and manual work Tags: leadership, sre, transformation ### Operations Team Lead — Workday Aug 2018 → Mar 2022 - Served as lead during major incidents, ensuring rapid resolution and clear communication - Streamlined incident response processes, reducing time-to-detect and time-to-resolve - Mentored junior engineers, elevating team technical and operational capabilities Tags: operations, incident-response, leadership ### Operations Engineer — Workday Nov 2017 → Aug 2018 - Triaged and resolved production incidents across customer-facing environments - Built Slack automations to reduce administrative toil and accelerate response - Contributed to deployment reliability through reviews and continuous improvements Tags: operations, automation, incident-response ### Network Operations Center Engineer — Robert Half Apr 2013 → Nov 2017 - Monitored and ensured network availability for field offices nationwide - Coordinated with ISPs and vendors to troubleshoot and resolve outages - Managed backup systems and performed datacenter tape rotations Tags: noc, networking, operations ## Projects and writing - **nerd.dad** (2024) — https://nerd.dad: Blog, homelab docs, and runbooks — SRE writing, GitOps manifests live elsewhere, prose lives here. - **Homelab GitOps Cluster** (2024) — https://nerd.dad/latest/homelab/: Talos Linux, Flux, Longhorn, and a full observability stack — production patterns at home scale. - **Hearth** (2025) — https://github.com/nerddotdad/hearth: Godot-powered 3D homelab visualizer with an agent bridge — because kubectl get pods isn't immersive enough. - **SRE Today 2025** (2025) — https://nerd.dad/latest/blog/posts/sre/sre_today_2025/: Notes from the field — what SRE looks like when AI, platform teams, and pager duty all show up to the same meeting. ## Skills - **Platform & Delivery:** Kubernetes, Helm, Jenkins, Git / GitHub Actions, CI/CD, Product Deployment, Bash / Shell Scripting - **Reliability & Observability:** Site Reliability Engineering, Incident Response, Grafana, BigPanda, Jira / JSM, 99.99% Availability Targets, Toil Reduction & Automation - **Leadership & Practice:** Engineering Management, Technical Product Management, Agile / Scrum, Mentorship, Python, Effective Communication, Critical Thinking ## Homelab Self-hosted Talos Kubernetes cluster, managed with Flux GitOps. Public docs: https://nerd.dad/latest/homelab/ - Talos Linux (Kubernetes) - Flux CD (GitOps) - Longhorn (Storage) - Observability (Prometheus · Grafana · Hearth) - Media (Jellyfin · Immich) - GitOps repo (truecharts / clusters/main) ## Links - GitHub: https://github.com/nerddotdad - nerd.dad: https://nerd.dad - LinkedIn: https://www.linkedin.com/in/tjhoth - Site: https://tjhoth.me - Support chat: https://support.tjhoth.me