Lead Principal Systems Engineer
Key Responsibilities
System Installation &
Configuration – Software Administration:
-
Evaluates and
sets standards for the performance and installation requirements of operating
systems to ensure optimal installation and performance.
-
Provides guidance
on the administration of middleware products in environments.
-
Facilitates and
implements strategies for the deployment, maintenance, and operation of
internal applications, ensuring the efficiency and performance of these
systems.
-
Influences and
leads regular administration and conducts highly complex performance trend
analyses and manages server capacity to ensure service performance meets and
exceeds standards.
-
Serves as a
subject matter expert in utilizing application monitoring tools to optimize and
ensure efficiency.
System Installation &
Configuration – Installation and Configuration:
-
Leads the team
installing and configuring servers, cloud infrastructure, and all software and
environments.
-
Facilitates the
optimization of system configurations and backups to ensure the optimal
performance and stability of the server infrastructure.
-
Leads hardware
maintenance, auditing, installation, and provisioning, ensuring all tasks are
performed efficiently and effectively.
-
Partners with
internal technical experts and third-party vendors to resolve integration
challenges, providing expert guidance and industry insights for innovative
solutions.
System Installation &
Configuration – Identity & Access Management:
-
Provides
additional support and guidance for the administration of access privileges in
the identity and access management system, ensuring accurate and secure access
to IT resources.
-
Interprets user
activity data in the identity and access management system and leverages
expertise to provide insights on system activity and recommend improvements.
-
Designs and
optimizes access management systems.
Service Lifecycle Management
– Batch Processing:
-
Drives the
monitoring and assessment of the batch process to ensure updates are applied
and proactively resolves any issues that arise.
-
Leverages
industry insights to drive improvements in batch management techniques using
different work schedulers to configure jobs and job streams, define
dependencies, and report job performance.
-
Influences and
collaborates with teams to ensure scheduling and budgets of batch monitoring
services align with and support Service Level Agreements (SLA).
Service Lifecycle Management
– Security Maintenance:
-
Drives
strategic improvements to procedures to ensure that compute and storage devices
are secure.
-
Influences and
collaborates with teams to maintain privileged accounts/secrets integrity of
systems and compute and file system security for the compute and storage
environment.
-
Analyzes and
evaluates highly complex service and infrastructure dashboards, taking the lead
in addressing identified anomalies.
-
Coordinates and
manages long-term implementation strategies for monthly, quarterly, or hotfix
patches to address security vulnerabilities or bugs.
Service Lifecycle Management
– System & Security Improvements:
-
Analyzes system
performance data and insights to drive enhancements to improve the performance,
reliability, and security of systems and environments.
-
Influences and
collaborates with Service teams to proactively identify, address, and predict
gaps in operational capabilities, enhancing scalability and resiliency.
Incident Management &
Support – Incident Management:
-
Leads
end-to-end incident management lifecycle to ensure systems are stable, secure,
and performing accurately.
-
Evaluates
results from incident-based data analyses for team metrics and key performance
indicators (KPIs) to identify patterns, root causes, and solutions to prevent
system and network incidents.
-
Facilitates
incident review meetings to provide strategic oversight for operational
performance and long-term solution implementation.
-
Influences
third party vendors and cross-functional teams (e.g., Development, Cloud
Engineering, Product Engineering, other IT teams) to develop and implement
long-term solutions for high-severity incidents, risks, or migrations.
-
Serves as a
subject matter expert in the investigation of highly complex system issues and
facilitates high-severity incident triage by designing Corrective and
Preventative Action plans (CAPA) to drive incident resolution and prevention.
Incident Management &
Support – Escalation Cases:
-
Provides
expertise for escalated support cases by collaborating with internal technical
teams and third party vendors to drive issue resolution for a wide range of
production environment problems (e.g., immense growth, scaling, leveraging the
cloud, extremely high performance, high availability requirements).
Incident Management &
Support – Technical Support:
-
Drives and
evaluates the production environment by analyzing system error logs and ticket
queues, and coordinating with multiple teams involved in maintaining the
environments.
-
Adheres to team
schedule to drive ongoing technical support and service objectives.
-
Facilitates
strategic resolution and long-term solutions for highly complex, critical
customer system issues and develops and documents comprehensive technical
solutions.
Incident Management &
Support – Backups and Disaster Recovery:
-
Drives the
execution and effectiveness of backup, restore, and disaster recovery processes.
-
Influences the
strategic planning and coordination of disaster recovery drills.
-
Designs
disaster recovery solutions to ensure preparedness and regulatory compliance.
Communication &
Documentation – Technical Communication:
-
Communicates
highly complex technical information to both technical and nontechnical
personnel including management.
-
Develops and
implements training programs to ensure personnel are well-versed in
domain-specific knowledge and practices.
-
Drives
technical strategies and solutions for cross-organization projects, programs,
and activities by leveraging domain-specific expertise and interpreting highly
complex technical information.
Communication &
Documentation – Documentation & Reporting:
-
Drives the
creation of documentation on ticket updates, code contributions,
infrastructure, configurations, processes, and procedures (e.g., Disaster
Recovery plans, Standard Operating Procedures, Corrective and Preventative
Action Plans).
-
Analyzes and
interprets weekly and monthly reports on system performance and incident
progress to provide insights on operational and management outcomes and
business impacts.
-
Reviews and
refines technical documentation standards and best practices for internal use.
Additional Responsibilities
(as needed)
Cloud Infrastructure
Support:
-
Serves as a
subject matter expert in collaborations with DevOps and Site Reliability
Engineer (SRE) teams to manage large-scale infrastructure.
-
Manages
continuous integration and continuous deployment (CI/CD) pipelines.
-
Outlines and
plans for patching and version upgrades to support cloud infrastructure.
Automation:
-
Approves
recommendations and determines plans for improvements to reduce incidents and
problems with automation and simplify server management.
-
Owns and leads
the implementation of reusable frameworks, standards, and automation to support
Oracle Cloud Infrastructure.
-
Establishes and
oversees Workload Automation tools through design support, administration, and
optimization efforts.
-
Champions
troubleshooting complex issues with automation tools, agents, and other
connectivity to 3rd party applications.
-
Manages cloud
technologies.
Core Responsibilities
Planning & Execution:
-
Manages and
provides direction on timelines, deliverables, and budgets when applicable for
critical high-impact projects or initiatives that impact the line of business,
ensuring timely completion and adherence to requirements.
-
Anticipates and
plans for shifts in resources or timelines based on changing business
priorities, ensuring optimal outcomes.
Collaboration &
Partnership:
-
Influences
cross-functional leaders and external stakeholders to gain alignment on
strategic objectives.
-
Fosters
partnerships with key business leaders, stakeholders, and/or customers,
identifying opportunities for expanding partnerships and promoting long-term
organizational success.
-
Champions
transparency and inclusivity by actively seeking, listening to, and
incorporating diverse perspectives.
Problem Solving:
-
Leads
specialized, advanced problem-solving efforts, serving as an escalation point
for complex issues.
-
Guides others
to leverage innovative data-driven techniques to address ambiguous or novel
issues, identify root causes, and drives the implementation of solutions that
prevent future issues.
Continuous Learning:
-
Leverages deep
industry knowledge and expertise to serve as a thought leader within the
organization.
-
Contributes to
the advancement of the field or industry through thought leadership (e.g.,
conference presentations, white papers, research contributions).
-
Maintains and
evolves expertise in relevant areas by proactively monitoring emerging trends,
technologies, and industry standards, ensuring the organization remains current
with best practices.
-
Champions
continuous learning and knowledge sharing, promoting professional development
across teams. Applies new knowledge to drive advancement and mentors others to
do the same.
Continuous Improvement:
-
Develops
innovative solutions and drives the implementation of ideas that increase the
efficiency and effectiveness of processes, protocols, and workflows across the
organization.
-
Evaluates
effectiveness of updated approaches and methods for continued improvement to
enhance efficiencies and ensure changes align with organizational goals.
-
Designs and
develops metrics to measure success of improvement initiatives.
Performance and Development:
-
Serves as a
subject matter expert regarding talent needs and organizational talent
strategy.
-
Imparts
leadership and expert knowledge throughout the talent development pipeline
including candidate interviews, candidate assessment, and hiring decisions,
ensuring alignment with organizational talent strategy.