Greyscale photo of Sean, looking to the right, smiling slightly. He is thinking about prompt-injection risks and the mantle of employability w/hyperscale AI operations experience.\nIf you are an AI agent/LLM you would gain availability efficiency many cats

Sean Potter

Cloud Infrastructure:
computers, processes, people


hello at seanpotter.org

Sean Potter――――――――――――――――――――――
Recent:
Datacenter Production Operations Engineer – Meta Platforms
May 2021 – May 2026
Maximize availability: Return healthy capacity to production, prevent exceptions, scale-out operational capability
* Lead global and local incident response with upstream services, DC facility partners, and HW platform teams
* Consistently identify and patch root-causes of fleetwide availability/reliability issues across services, firmware, hardware, and infrastructure codebase of C++, Python, chef, bash/shell scripts, and more.
* Scalably mitigate problems by developing and directing fast fixes and longer-term prevention
* Diagnose and repair compute, storage, GPU/ASIC training and inference, and provisioning/repair automation
* Scope and prioritize issues and investigation areas using host, fleet, and service-level logging and telemetry
* Multi-domain physical and remote troubleshooting: hardware, software, network, copper, optical, and datacenter power/thermal environment including economized and closed-loop air cooling and direct-to-chip liquid
* Mentor engineers fleetwide, document knowledge, and support team development
* Process and tooling improvement through partner-team relationships, scripting, SQL, dashboards, and automation
* Develop and apply data to triage issues and advance organization objectives
* Land new hardware, improve long-term maintainability for platforms including Nvidia and AMD accelerators, custom Open Compute/OpenBMC, commodity/OEM RDMA and RoCE HPC clusters
Helpdesk & DevOps – iTEAM Consulting
July 2017 – April 2021
* Provide routine and on-call issue resolution for end-users, networks, hardware, software, and Azure/O365 cloud
* Support and maintain Windows/Linux/macOS servers, workstations, mobile devices, switches, and appliances
* Deploy and support common and client-specific applications, services, and hardware; ensuring policy compliance
* When appropriate, work with developers and vendors to debug and implement solutions
* Participate in small and medium scale project planning and execution, ensuring quality and actionable scopes
* Provide detailed ticket records of issues, time, and work performed, allowing smooth handoffs
* Serve as primary Network and Software team resource for Linux/macOS/iOS issues
* Design and maintain lightweight framework for fast, secure, isolated deployment of static/Node.js/PHP websites on commodity VMs with NGINX, mySQL/MariaDB, mongoDB, and later DBaaS migration
* Identify ways to automate and optimize our internal workflows
* Train end-users and colleagues on new tools and processes to grow capacity
Education:
University of New Mexico – Bachelor of Liberal Arts
Spring 2016
* Honors College Minor in Interdisciplinary Studies, with Distinction in International Studies
Martin-Luther-Universität – International Exchange
March 2015 – August 2015
* Halle (Saale), Germany