Senior Operations Analyst (SRE)

Back

Senior Operations Analyst (SRE)

@ CGI

Position Description:

CGI's Advantage Cloud Operations is an SRE-driven operating model, anchored on an Operations Control Plane that unifies telemetry, event management, automation, and IT Service Management (ITSM). The Senior Operations Analyst is a senior Site Reliability Engineering (SRE) practitioner responsible for driving reliability engineering, problem management, and runbook automation across customer environments. This role also mentors the operations team in adopting proactive, data-driven operational practices while improving platform reliability, operational efficiency, and service quality.

Your future duties and responsibilities:

Reliability Engineering & Operations Leadership

    • Serve as a senior SRE within the Cloud Operations team supporting the CGI Advantage platform.
    • Own end-to-end reliability outcomes, including reducing MTTD and MTTR, minimizing escalations, and improving release safety.
    • Lead incident response, Root Cause Analysis (RCA), and Problem Management activities.
    • Design and implement automation and self-healing workflows with appropriate guardrails, approvals, and rollback capabilities.
    • Act as the technical lead for the Operations technology stack and drive platform adoption.
    • Mentor Operations Analysts while establishing operational standards, taxonomy, and Service Level Objective (SLO) discipline.

Observability & Incident Response

    • Drive distributed tracing adoption using OpenTelemetry with trace-to-log and metric correlation for priority services.
    • Define SLO-aligned alerting strategies and optimize deduplication, suppression, and event correlation to reduce alert noise.
    • Lead Sev 1 and Sev 2 incident bridges while ensuring evidence-based RCAs and corrective actions are completed.
    • Standardize dashboards and observability signal sets across customer environments.

Automation & Problem Management

    • Build and maintain an active automation backlog and deliver automated runbooks on a quarterly basis.
    • Convert recurring incidents into problem records with RCA hypotheses, ownership, and due dates.
    • Develop guard-railed self-service operational actions (diagnostics, restarts, environment tasks) utilizing RBAC and audit trails.
    • Champion Configuration-as-Code practices, including version control, drift detection, and baseline versus override governance.

Release Safety & Security

    • Enforce release gates, quality thresholds, and automated rollback to the last known good state.
    • Implement time-bound privileged access (JIT/Break Glass), SSO integration, and audit-ready operational controls.
    • Apply Policy-as-Code guardrails across Infrastructure as Code (IaC), Kubernetes admission controls, and CI/CD pipelines.

Leadership & Governance

    • Mentor and upskill Operations Analysts while documenting and promoting "golden path" operational workflows.
    • Report reliability KPIs including: Mean Time to Detect (MTTD) , Mean Time to Resolve (MTTR) , First Touch Resolution , Automation Coverage
    • Partner with Product, Delivery, ITSM, and Security teams within a matrix operating model.


Qualifications:

Required qualifications to be successful in this role:

  • 6–9 years in cloud operations/SRE, including 3+ years operating production SaaS on Azure.
  • Observability: OpenTelemetry, distributed tracing, log analytics, metrics, and dashboards.
  • Kubernetes at scale: AKS, Calico network policy, KEDA autoscaling, Helm based deployments.
  • IaC and configuration: Terraform; GitOps with GitHub Actions and Argo CD.
  • Event management and orchestration: event correlation, auto remediation workflows.
  • Scripting/automation: strong Python and Bash; API driven integration across ITSM and tooling.
  • Security tooling: Keycloak/ SSO, Azure Key Vault, cert manager, JIT access patterns.
  • ITSM depth: incident/problem/change management, severity taxonomy, SLA/SLO reporting.

Preferred Experience

  • Policy as code (OPA, Kyverno, Conftest) and compliance grade audit reporting.
  • Job/batch orchestration (JS7 or equivalent) and API gateway operations (APISIX/APIM).
  • PostgreSQL operations, PgBouncer connection pooling, and managed database patterns.
  • Multi tenant managed services delivery with contractual SLA/SLO obligations.
  • Certifications: CKA/CKAD, Azure Administrator/Architect, ITIL 4.


How to Apply:

Apply online at: https://www.cgi.com/en/careers 

Visit Site to Apply

Location: Lafayette, LA
Date Posted: July 30, 2026
Application Deadline: August 31, 2026
Job Type: Full-time