Senior Linux Network Engineer
Role Overview:
Itero Group is seeking a Senior Linux & Network Engineer with strong hands-on troubleshooting skills across enterprise Linux systems, IP networks, and highly available production infrastructure.
Key Responsibility Areas
Act as a first technical point of response and primary customer technical contact during production incidents, owning the investigation from initial detection through service restoration and follow-up.
Perform hands-on troubleshooting across the complete service environment, including Linux systems, IP networks, firewalls, Pacemaker/Corosync clusters, Galera database clusters, virtualization, applications, and customer connectivity.
Diagnose complex Linux and network problems using command-line tools, logs, monitoring data, routing information, and packet captures using tools such as tcpdump and Wireshark.
Troubleshoot routing, switching, firewall, VPN, TCP/UDP, performance, and connectivity problems, including packet loss, latency, asymmetric routing, and other service-impacting network conditions.
Troubleshoot highly available Linux and database environments, including Pacemaker/Corosync and Galera, with issues involving quorum, failover, cluster communication, replication, synchronization, and recovery.
Coordinate internal engineering teams, customers, network/data-center providers, and other third parties when additional assistance is required while maintaining technical ownership of the incident.
Design, deploy, maintain, patch, harden, monitor, and improve Etherstack's Linux, network, clustering, and cloud infrastructure.
Perform root-cause analysis, execute controlled production changes, maintain technical documentation, and develop Bash/Python automation to improve operations and troubleshooting.
Required Experience and Skills
This position requires an experienced engineer who can contribute independently in a production environment from the beginning of employment. Candidates must already possess strong hands-on Linux, networking, and troubleshooting experience. Candidates will be introduced to Etherstack-specific applications, architecture, and operational procedures; however, this role is not intended to provide foundational training in Linux systems administration, networking, high availability, or production incident response.
Linux: Strong hands-on experience administering and troubleshooting production Linux environments, preferably Red Hat Enterprise Linux (RHEL). Candidates should be comfortable diagnosing services, processes, CPU/memory, storage and filesystems, networking, security, performance, and application-connectivity issues from the command line.
Networking: Strong knowledge of TCP/IP, routing, and switching, with hands-on experience troubleshooting VLANs, routing protocols such as OSPF and BGP, firewalls, NAT, VPNs, TCP/UDP connectivity, packet loss, latency, and routing problems. Experience with Juniper and/or Cisco infrastructure is preferred.
High Availability: Hands-on experience supporting and troubleshooting highly available or clustered production systems, including concepts such as quorum, failover, node communication, fencing, replication, synchronization, and recovery. Experience with Pacemaker/Corosync and Galera Cluster is strongly preferred.
Troubleshooting: Demonstrated ability to investigate complex problems across multiple technology layers using Linux diagnostic tools, logs, monitoring systems, routing information, and packet captures with tcpdump and/or Wireshark. Candidates should be able to methodically isolate whether a problem originates in the operating system, network, database, cluster, application, or an external dependency.
Production Operations: Experience supporting highly available, business-critical or mission-critical production environments, including incident response, change management, root-cause analysis, and direct technical communication with customers during service-impacting incidents.
Automation: Working knowledge of Bash and/or Python for Linux administration, troubleshooting, and operational automation.
Candidates typically have 5+ years of hands-on experience supporting production Linux systems and IP networks, including responsibility for troubleshooting service-impacting incidents.