Lead Cloud Performance Engineer (A&D, Ultra HA/Exadata)
- IFS
- Colombo, Sri Lanka
- LKR 6,000,000 – LKR 9,000,000
Company Description
IFS is a billion-dollar revenue company with 7000+ employees on all continents. Our leading AI technology is the backbone of our award-winning enterprise software solutions, enabling our customers to be their best when it really matters–at the Moment of Service™. Our commitment to internal AI adoption has allowed us to stay at the forefront of technological advancements, ensuring our colleagues can unlock their creativity and productivity, and our solutions are always cutting-edge.
At IFS, we’re flexible, we’re innovative, and we’re focused not only on how we can engage with our customers but on how we can make a real change and have a worldwide impact. We help solve some of society’s greatest challenges, fostering a better future through our agility, collaboration, and trust.
We celebrate diversity and understand our responsibility to reflect the diverse world we work in. We are committed to promoting an inclusive workforce that fully represents the many different cultures, backgrounds, and viewpoints of our customers, our partners, and our communities. As a truly international company serving people from around the globe, we realize that our success is tantamount to the respect we have for those different points of view.
By joining our team, you will have the opportunity to be part of a global, diverse environment; you will be joining a winning team with a commitment to sustainability; and a company where we get things done so that you can make a positive impact on the world.
We’re looking for innovative and original thinkers to work in an environment where you can #MakeYourMoment so that we can help others make theirs. With the power of our AI-driven solutions, we empower our team to change the status quo and make a real difference.
Job Description
The role of IFS Lead Cloud Performance Engineer exists within the Unified Support organization/division. As part of a shift operation, they will contribute towards providing 24x7x365 support to IFS customers across the globe within the Aerospace & Defense (A&D) industry vertical.
This position sits within the Unified Support Performance Team dedicated to our Ultra High Availability (Ultra HA) on Oracle Exadata customers. The team owns proactive performance monitoring, deep-dive diagnostics, tuning and capacity management across the full stack Oracle Database on Exadata, the IFS application tiers, and the supporting cloud and network infrastructure to keep these mission-critical environments fast, stable and continuously available.
As Lead Cloud Performance Engineer, you will be responsible for handling performance-related technical issues reported by customers, service requests, and problem and change management. You will lead performance investigations end to end from detection in the monitoring and observability platforms, through root cause analysis on the Oracle Database and application tiers, to the permanent fix or tuning recommendation. This role involves collaborating with various internal and external stakeholders to ensure problem-solution fits, which ultimately leads toward customer success. You will also play a crucial role in development and maintenance of IFS's products and internal systems.
- Work with other Service Center functions and appropriate stakeholders to resolve long running, complex or major incidents.
- Create and update relevant SOPs, FAQs and other documentation to address known issues, workarounds, and service requests.
- Manage an incoming queue of cases, incidents, and service requests within SLA, OLA and KPI targets.
- Support the event management team and their work to enhance the related event processes and tools.
- Proactively monitor the performance and availability of Ultra HA on Exadata environments, and detect, triage and resolve degradation before it impacts the customer's Moment of Service.
- Perform deep-dive Oracle Database performance troubleshooting wait event and AWR/ASH analysis, SQL and execution plan tuning, RAC and Data Guard behavior, storage and I/O bottlenecks on Exadata.
- Analyze end-to-end application performance across the IFS application tiers using APM and observability tooling, correlating application traces with database and infrastructure telemetry.
- Build and maintain dashboards, alerts and SLO/SLI reporting in Elasticsearch, Grafana and Open Telemetry-based tooling, and continuously reduce alert noise and false positives.
- Lead performance root cause analysis for major incidents and produce tuning, configuration and capacity recommendations for customers and internal stakeholders.
- Contribute to capacity planning, load and stress testing, and pre-upgrade or pre-release performance validation for Ultra HA customers.
- Boost productivity significantly with AI by reducing redundant, admin‑heavy tasks and enabling teams to focus on higher‑value work.
- Liaise with IFS R&D and Unified Support Engineering organizations for automation requirements to support cloud-based applications.
- Liaise with IFS R&D and Unified Support Engineering organizations for knowledge sharing with Unified Support Service Operations groups in relation to cloud-based applications.
- Identifying knowledge gaps and compile/ update internal Knowledge Based Articles (KBA) to guide Unified Support Service Operations groups in relation to supporting cloud-based applications.
- Identify and communicate improvement points related to Standard Operating Procedures (SOP) and Processes to the relevant process owners.
Qualifications
- A university degree or an equivalent professional qualification in Software Engineering, Computer Science, Information Technology or similar.
- Experience in a modern ticket/service desk tooling such as ServiceNow, Jira Service Desk, or a similar tool.
- Experience in ITIL, ISO 20000, or a similar IT service delivery framework.
Mandatory Skills
- At least 7+ years’ experience in cloud computing services, enterprise IT delivery services or similar, including demonstrable experience troubleshooting performance in production Oracle Database and enterprise application environments.
- Understanding the low-level concepts in Cloud computing.
Oracle Database performance monitoring and troubleshooting, together with application performance monitoring, are essential for this role. The successful candidate will additionally have at least half of the remaining skills below, or a suitable professional grade qualification.
- Cloud service administration and operations (Azure, AWS & GCP).
- Demonstrate proficiency in debugging complex applications.
- Linux Server administration.
- Network administration.
- Oracle Database administration, with strong monitoring, performance diagnostics and troubleshooting skills (AWR/ASH/Statspack, SQL and execution plan tuning, wait event analysis, Oracle Enterprise Manager).
- Working knowledge of Oracle Exadata and high-availability architectures (RAC, Data Guard, ASM), including their performance characteristics.
- Application performance monitoring and observability instrumenting, analyzing and tuning application-tier performance.
- Hands-on experience with Elasticsearch, Grafana, Open Telemetry or comparable APM/observability tooling (for example Dynatrace, AppDynamics, New Relic, Prometheus).
- Web Server administration (Wildfly, WebLogic & Nginx).
- Kubernetes/Docker operations and administration.
- BASH/PowerShell/Terraform/Ansible scripting usage.
Soft Skills
- Ability to work to deadlines and targets.
- Ability to manage own time efficiently and effectively.
- Ability to work in international, multi-disciplined, cross-functional teams.
- Flexibility to work to deadlines and needs of the role.
- Ability to read and understand technical documentation written in English.
- Ability to mentor and provide a good role model for junior team members.
- Problem-solving skills and the ability to change approach based on information gathered during the process.
- Good communication and interpersonal skills.
- Strong organizational skills and ability to multi-task.
- A positive team player with a can-do attitude.
- Excellent verbal and written communication skills in English.
- Ability to self-learn and quickly understand new and changing technologies in a fast-moving service driven technology landscape.
- Proactivity and ownership of work items in all aspects of the technical and team role.
Additional Information
We embrace flexibility and hybrid work opportunities to support diverse needs and lifestyles, while also valuing inclusive workplace experiences. By fostering a sense of community, we drive innovation, strengthen connections, and nurture belonging. Our commitment ensures you can work in a way that suits you best, while also engaging with colleagues to share ideas and build meaningful relationships.
Skills
- Cloud Performance Engineering
- Oracle Exadata
- High Availability Architecture
- Performance Tuning
- Load Testing
- Linux Systems Administration
- IT infrastructure monitoring







