We use cookies. Find out more about it here. By continuing to browse this site you are agreeing to our use of cookies.
#alert
Back to search results
Remote

Principal Database Engineer

NextGen Healthcare
United States, Georgia
Sep 14, 2026

Job Description:

The Principal Database Engineer (DBE) - Cloud Platform Support is a premier technical authority and operational leader within the Cloud Platform Support organization. This role bridges deep database engineering with customer-centric support workflows and best practices, close collaboration with Site Reliability Engineering (SRE) and DevOps practices to guarantee the availability, reliability, and peak performance of NextGen's enterprise database platforms supporting cloud-first healthcare applications.
Operating at the intersection of database architecture and complex operational support, this engineer serves as the primary Subject Matter Expert (SME) for Tier-3 database escalations, query and engine performance optimization, and operational telemetry. They are instrumental in eliminating operational toil through self-healing automation, defining operational readiness standards for cloud database migrations, and mentoring support engineers to continuously elevate frontline technical capabilities.

Key Responsibilities

Technical Leadership & Support SME Authority


  • Serve as the dedicated Subject Matter Expert (SME) for the Support organization on database engine internals, performance troubleshooting, concurrency bottlenecks, and functional architecture across Microsoft SQL Server, PostgreSQL, Linux and cloud-native databases.
  • Act as the final technical escalation point for critical production database incidents (Sev-1/Sev-2), leading immediate triage, dynamic query tuning, index restructuring, and system recovery.
  • Build, standardize, and maintain comprehensive database troubleshooting runbooks, diagnostic playbooks, and self-service diagnostics for frontline teams.
  • Conduct regular technical enablement sessions, knowledge transfers, and mentoring programs to upskill Support engineers in advanced database telemetry, log analysis, and performance triage.

SRE, DevOps & Support Process Integration


  • Bridge the gap between Core Product Engineering, Database Platform/DevOps teams, and Support operations, ensuring production-support feedback directly influences upstream architectural roadmaps and deployment strategies.
  • Partner with Platform and Release Engineering to establish robust database release governance, automated schema validation, and safe zero-downtime deployment pipelines that prevent support regressions.
  • Embed established Site Reliability Engineering (SRE) principles into daily support operations, actively defining and tracking database Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
  • Champion operational toil reduction by designing and deploying event-driven automation, auto-remediation scripts, and self-healing workflows for recurring database health events.

Platform Observability, Telemetry & Health Monitoring


  • Architect and operationalize end-to-end database observability standards across public cloud environments (AWS, GCP), instrumenting granular telemetry, query performance analytics, wait-state profiling, and proactive alerting.
  • Leverage advanced APM and observability platforms (e.g., Dynatrace, Prometheus/Grafana, native cloud monitoring) to identify latent performance degradation before it impacts client experience.
  • Define and maintain the database tier of the overall Cloud Platform Health Score, providing continuous visibility into capacity limits, storage growth, replication latency, and connection saturation.

Incident Remediation, Post-Mortems & Operational Excellence


  • Drive blameless post-incident reviews (PIRs) and Root Cause Analysis (RCA) investigations for major client-facing database incidents, tracking corrective and preventive actions (CAPA) through to completion.
  • Analyze recurring support ticket trends, escalation patterns, and Mean Time to Resolution (MTTR) metrics to identify systemic database flaws and drive engineering remediation backlogs.
  • Formulate robust business continuity (BC) and disaster recovery (DR) protocols, leading scheduled validation exercises for automated failovers, cross-region replication, and point-in-time recovery (PITR).

Database Modernization & Operational Readiness


  • Act as the operational readiness gatekeeper for enterprise migration initiatives transitioning legacy Microsoft SQL Server workloads toward PostgreSQL and cloud-native database services.
  • Coordinate operational handoffs and Hyper Care support windows for newly migrated workloads, validating that monitoring, backup policies, and rollback mechanisms meet stringent enterprise SLA standards.
  • Continuously evaluate data-tier technologies and operational tooling to recommend architectural adjustments that enhance platform resiliency, maintainability, and client satisfaction.

Qualifications & Experience

Education


  • Bachelor's Degree in Computer Science, Information Technology, Software Engineering, or a related technical field; or an equivalent combination of education, technical certifications, and extensive professional experience.

Required Experience


  • 10+ years of hands-on experience designing, supporting, managing, and tuning enterprise relational database platforms in high-volume, mission-critical 24/7 production environments.
  • 4+ years operating enterprise databases within public cloud platforms (AWS, GCP, or Azure), utilizing Infrastructure as Code (Terraform, CloudFormation) and CI/CD pipelines.
  • Demonstrated track record serving as a senior technical authority or principal SME within an Operational Support, SRE, or Production Engineering organization handling complex escalations and client-facing incidents.
  • Proven mastery of Microsoft SQL Server internals (indexing, execution plans, locking/blocking, AlwaysOn AGs, tempdb contention) alongside strong operational experience with PostgreSQL, SQL over Linux and cloud-native relational engines (Amazon Aurora, Cloud SQL).
  • Extensive experience conducting deep root cause analysis (RCA), writing technical post-mortems, and driving structural remediations to reduce incident volume and MTTR.
  • Experience implementing proactive database observability frameworks using modern APM/telemetry tooling (e.g., Dynatrace, Datadog, Prometheus, Grafana).
  • Background supporting distributed, cloud-native microservices or Kubernetes-based workloads within SaaS, healthcare, or other highly regulated environments (preferred).

Preferred Certifications


  • Microsoft Certified: Azure Database Administrator Associate or legacy MCSE/MCSA SQL Server.
  • AWS Certified Solutions Architect - Associate/Professional or AWS Certified Database - Specialty.
  • Google Cloud Professional Cloud Database Engineer.
  • Certified Kubernetes Administrator (CKA).
  • HashiCorp Certified: Terraform Associate.
  • ITIL Foundation or SRE Practitioner certification.

Knowledge, Skills & Abilities

Technical Knowledge


  • Database Internals & Administration: Expert-level understanding of Microsoft SQL Server and PostgreSQL engine mechanics, memory management, storage subsystem I/O patterns, buffer cache, locking/concurrency models, and HADR topologies.
  • Modern Cloud Architecture: Deep familiarity with AWS/GCP managed database architectures, containerized workloads, security compliance (HIPAA, HITRUST, SOC2), encryption at rest/in transit, and role-based access control (RBAC).
  • Observability & Analytics: In-depth knowledge of database telemetry instrumentation, wait-event analysis, distributed tracing, and metrics aggregation using platforms like Dynatrace.

Core Skills


  • Performance Tuning & Diagnostics: World-class ability to diagnose complex, multi-threaded database performance bottlenecks, decipher complicated query plans, optimize dynamic T-SQL/PL-pgSQL, and eliminate deadlocks.
  • Automation & Scripting: Strong scripting and programming skills (Python, PowerShell, Bash, SQL) to build automated diagnostics, health checks, and self-healing runbooks that replace manual toil.
  • Process & Communication: Exceptional capacity to translate deep technical findings into concise, executive-ready incident summaries, client RCA reports, and actionable development tickets.

Professional Competencies


  • Cross-Functional Bridge: Ability to seamlessly collaborate with developers, cloud engineers, product managers, support managers, and client-facing teams to resolve complex platform challenges.
  • High-Pressure Composure: Demonstrated poise and decisive technical leadership during high-severity production outages.
  • Mentorship & Culture: Passion for coaching, developing frontline support engineers, and fostering an operational culture centered on blameless post-mortems, automation, and continuous improvement.

The company has reviewed this job description to ensure that essential functions and basic duties have been included. It is intended to provide guidelines for job expectations and the employee's ability to perform the position described. It is not intended to be construed as an exhaustive list of all functions, responsibilities, skills and abilities. Additional functions and requirements may be assigned by supervisors as deemed appropriate. This document does not represent a contract of employment, and the company reserves the right to change this job description and/or assign tasks for the employee to perform, as the company may deem appropriate.

NextGen Healthcare is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Applied = 0

(web-9db6c7984-zd648)