Anthropic

Incident Management Lead, Data Center Security

Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NYCompany updated 1 hour agoVerified 8/27/2026
Security
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role Anthropic operates and is building data center campuses around the world, run day to day through operating partners, site vendors, and contracted security services. Things go wrong at those sites the way they go wrong at any critical infrastructure: access issues, contractor incidents, protests, weather, equipment failures, and occasionally something with the potential for global impact. Today, what counts as an incident, who gets told, and how it is reported vary by site and by vendor. As Lead for Crisis and Incident Management, you will build the program that removes that variance, and you will build it with the people who have to live with it. The first job is definition: working with operating partners, site security vendors, managed-service providers, and internal teams to agree on what counts as a minor escalation versus a major incident of global impact, the severity levels in between, the thresholds that move an event from one level to the next, and who gets told, how fast, and in what format. The second job is adoption: turning those definitions into playbooks, training, and exercises so that responders at every site and vendor actually respond the right way, and running the after-action reviews that show where they did not. The third job is measurement: defining the incident metrics, building the reporting, and giving leadership a regular, accurate picture of how response is performing across the fleet. Underneath all three sits the cross-functional work of getting partners to buy in, because a framework no partner has signed up to is a document, not a program. Continuous coverage is part of the picture. As the program matures, you will shape a 24/7 monitoring and response capability through GSOC-type managed services and site vendors, so that the framework holds at three in the morning at any site in the fleet. That build-out follows from the definitions and partner agreements above rather than replacing them. When a major incident does happen, you run it: activation, coordination, leadership communication, stand-down, and the after-action review that makes the program better. This is a program-building role, not a shift-supervision role. The deliverable is a framework that partners have adopted, that works at any site, and that reports on itself, run largely through vendors and managed services rather than a large in-house team. Where existing processes and vendor arrangements don't support that, you will have the authority, and the expectation, to change them. Role boundaries. This seat is distinct from the Data Center Security Delivery Lead (who owns security through construction, commissioning, and handover on new builds), from regional security operations leads (who own steady-state regional programs and vendor relationships site by site), and from the team's systems-engineering roles (who build platforms and tooling). This role owns the horizontal incident and crisis layer across the operating fleet: the incident definitions and severity framework, partner adoption and governance, incident reporting and metrics, major incident command, post-incident review, and, as it develops, the 24/7 coverage model, applied consistently across every site and vendor. Key responsibilities You will build and own the global crisis and incident management program for data center physical security. Incident definition and severity framework: define, with partners, what counts as an incident and the tiers from minor site escalation to major incident of global impact, with clear thresholds, escalation and activation criteria, notification requirements, and decision rights at each tier. This is the foundation the rest of the program is built on. Cross-functional partner buy-in and governance: bring operating partners, site security vendors, managed-service providers, and internal teams into the definition work so the framework is theirs as well as ours, and establish the governance (owners, review cadence, change process) that keeps it current as the fleet grows. Adoption and response readiness: turn the framework into playbooks, training, and tabletop and functional exercises across sites and vendors, so responders know the right response before the next real incident and the escalation chain is tested by design. Metrics and reporting: define the incident metrics that show whether definitions are being applied and whether response is improving, build the dashboards and reporting behind them, and own the executive reporting cadence, from real-time notification through leadership escalation and executive summaries. Vendor and managed-service consistency: hold multiple vendors and GSOC-type providers to one standard: common definitions, common procedures, common reporting formats, common escalation triggers, with performance measured and enforced. Major incident management: run the crisis process when an incident has multi-site or global impact: activation, coordination across sites, vendors, and internal teams, decision support to leadership, and formal stand-down. Post-incident review and improvement: run structured after-action reviews, track corrective actions to closure, and feed lessons back into definitions, playbooks, procedures, and vendor contracts. 24/7 coverage model: as the program matures, design and run an always-on monitoring, escalation, and response capability through GSOC-type managed services and site security vendors, achieving continuous coverage without building a large in-house shift operation. Minimum qualifications Have built or substantially rebuilt an incident management or crisis management program (not just operated within one), including the definitions and severity model at its core, and can describe what existed before you, what you changed, and how you know it worked. Have brought partners you don't control (operators, vendors, internal teams) into agreeing on definitions and procedures, and can describe how you won that buy-in and kept it. Have driven adoption: you have taken a framework from paper into playbooks, training, and exercises, and can point to responders behaving differently as a result. Have defined incident metrics and built the reporting behind them, and can explain which measures actually told leadership something and which did not. Know data centers: you have worked physical security in or around data center or comparable critical-infrastructure operations, and understand the environment of operating partners, contractors, and 24/7 site activity. Have run major incidents end to end: activation, multi-party coordination, leadership communication, stand-down, and after-action, and can walk through specific ones. Have held vendors or managed services (a GSOC, monitoring provider, or guard force) to defined performance standards, with evidence rather than assurances. Can make the severity call quickly on incomplete information, defend it either way, and adjust as facts arrive. Write clearly under pressure: your incident report-outs can go to executives without editing. Work through influence across sites, vendors, and internal teams you don't control; you shape response rather than waiting to be given authority. Preferred qualification Experience designing incident taxonomies, severity models, or escalation frameworks that were adopted across multiple organizations or vendors. Experience running an incident metrics program at scale (dashboards and an executive reporting cadence spanning multiple sites or vendors). Incident command system experience (ICS/NIMS or comparable). Experience standing up or running a global security operations center (GSOC) or equivalent 24/7 capability. CPP (Certified Protection Professional), PSP (Physical Security Professional), CEM (Certified Emergency Manager), CBCP, or comparable certifications. Hyperscaler, major colocation, or critical-infrastructure operator experience. Business continuity or emergency management program background. Experience in regulated or high-assurance environments. The annual compensation range for this role is listed below. For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role. Annual Salary: $290,000—$365,000 USD Logistics Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices. Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this. We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team. Your safety matters to us. To protect yourself from potential scams, remember that Anthropic recruiters only contact you from @anthropic.com email addresses. In some cases, we may partner with vetted recruiting agencies who will identify themselves as working on behalf of Anthropic. Be cautious of emails from other domains. Legitimate Anthropic recruiters will never ask for money, fees, or banking information before your first day. If you're ever unsure about a communication, don't click any links—visit anthropic.com/careers directly for confirmed position openings. How we're different We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. And we value impact — advancing our long-term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time. As such, we greatly value communication skills. The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences. Come work with us! Anthropic is a public benefit corporation headquartered in San Francisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues. Guidance on Candidates' AI Usage: Learn about our policy for using AI in our application process.