DevOps & Infrastructure Engineer (On-Premises)
--Fortek Pvt Ltd.--
Job Description:
Fortek (Private) Limited is expanding its
on-premises ICT infrastructure and data center footprint, including GPU-enabled
compute for AI/ML workloads. The DevOps & Infrastructure Engineer will be
the operational backbone of this environment — owning the full lifecycle of
company-owned physical and virtual servers, from provisioning and hardening
through to automation, monitoring, and disaster recovery. This is a hands-on,
individual-contributor role for someone who is equally comfortable in a server
room, a Kubernetes cluster, and a runbook. The incumbent will work closely with the Head of
Engineering/CTO and cross-functional engineering teams to build repeatable,
secure, and well-documented infrastructure practices, while directly supporting
local production systems, CI/CD pipelines, and GPU-accelerated deployments that
underpin Fortek's service delivery to clients across Pakistan's ICT and data
center sector.
Position Structure:
Department:
Line Manager:
Stream:
Job Requirements
-
Skills & Tools:
Educational Qualification:
- Bachelor's degree in Software Engineering, Computer Science, Information Technology, or a related field.
Certifications (Preferred)
- RHCSA/RHCE, CKA/CKAD, NVIDIA-Certified Associate (GPU/CUDA), CompTIA Linux+/Network+, or equivalent vendor certifications.
Other Skills & Tools
- Linux administration; bare-metal and virtualized infrastructure (VMware, Proxmox, Hyper-V); Docker/Kubernetes; Ansible; Git; Nginx/HAProxy.
- Grafana, Prometheus, Loki, Tempo/OpenTelemetry; networking fundamentals (VLANs, DNS, DHCP, VPNs, firewalls).
- NVIDIA GPU servers, drivers, CUDA and container toolkit; storage, backups, access control, patching and disaster recovery.
Job Experience Required:
- 2+ years of hands-on experience in DevOps, systems administration, infrastructure engineering, platform engineering, or a related role.
- Experience supporting local production servers, building CI/CD pipelines, automating infrastructure and troubleshooting live systems.
- Experience deploying applications or AI/ML workloads on GPU-enabled servers, including GPU containers and monitoring, is a strong advantage.
Perks & Benefits:
At Fortek, we believe in empowering our people. We offer:
Competitive salary based on experience and qualifications
Annual performance-based increments & bonuses
Medical facility
Paid annual, casual & sick leaves
Gratuity and EOBI benefits
Training & professional development opportunities
Collaborative, learning-driven work environment, and much more
Let’s build the future together! ⚡
Duties & Responsibilities
-
1. Infrastructure Design, Installation & Maintenance
- Design, install, configure and maintain secure production infrastructure on company-owned physical and virtual servers.
- Administer Linux systems and virtualization clusters (VMware, Proxmox or Hyper-V); manage compute, memory, storage, OS patching and capacity.
- Maintain working knowledge of Windows Server for mixed environments.
2. DevOps Automation & CI/CD
- Build and operate self-hosted CI/CD runners, source-control integrations, artifact repositories and container registries.
- Automate provisioning, configuration and repeatable operations using Ansible, scripting (Bash/Python) and infrastructure-as-code practices.
- Manage Docker, Docker Compose and Kubernetes deployments, including GPU-aware scheduling.
3. GPU & AI/ML Workload Enablement- Deploy containerized applications and AI/ML workloads on GPU-enabled local servers using Docker/Kubernetes.
- Maintain NVIDIA GPU drivers, CUDA and container runtimes; monitor GPU utilization, memory, temperature, health and workload performance.
4. Network & Security Administration- Manage internal networking: VLANs, routing, DNS, DHCP, VPNs, reverse proxies, load balancers (Nginx/HAProxy) and firewalls.
- Apply least-privilege access controls, secrets management, hardening, vulnerability remediation and audit logging.
- Maintain TLS certificate lifecycle across internal and client-facing services.
5. Monitoring, Backup & Disaster Recovery
- Implement and maintain observability using Grafana, Prometheus, Loki and Tempo/OpenTelemetry, with centralized logging.
- Own backups, restore testing, replication, high availability, incident response, root-cause analysis and disaster recovery.
6. Documentation & Stakeholder Coordination
- Maintain inventory, architecture diagrams, SOPs and runbooks.
- Coordinate infrastructure changes with developers and provide on-call support as required.
- Log all incidents, service requests, risks and escalations in Odoo per Fortek's agreed tracking protocol.
Reporting Responsibilities (Daily, Weekly & Monthly)
-
Daily Reporting:
Daily Tasks:
• Monitor server, GPU, storage, network, application, database, backup, deployment and pipeline health; respond to alerts and incidents.
• Resolve incidents and service requests; log actions, risks and escalations in Odoo Helpdesk/Task modules.
Weekly Reporting:
Weekly Tasks:
• Submit deployment, incident, reliability and capacity summaries to the Line Manager and engineering stakeholders via Odoo.
• Review infrastructure changes, patching, backup/restore status, vulnerabilities, access and capacity items.
Monthly Reporting:
Monthly Tasks:
• Submit KPI report in Odoo covering uptime, incidents, MTTR, deployment performance, resource utilization, backup success and security posture.
• Review CPU/GPU/server/storage capacity, lifecycle needs, DR readiness, documentation and the infrastructure roadmap.