Production Engineer, Finance Tech
Balyasny Asset Management LP Warsaw, PolandProduction Engineer, Finance Tech
Balyasny Asset Management is a global institutional investment firm where finance and technology work together to support investment teams, manage risk, and deliver superior service to our investment partners. With more than 2,000 [JS1] employees across offices worldwide, BAM offers a collaborative, high-performance environment built around innovation, continuous learning, and professional growth.
Role Overview
Are you an engineer who enjoys solving complex production problems as much as building the solutions that prevent them?
As a Production Engineer, you will combine hands-on engineering with ownership of live production systems, working closely with traders, portfolio managers, operations teams, and technology teams.
The role is split between live support and engineering projects. During support rotations, you will play a leading role in our global follow-the-sun support model, monitoring production and providing first- and second-line support for applications spanning trading, P&L, accounting, treasury, and related business functions.
This includes diagnosing and resolving live issues, coordinating the right technical and business stakeholders, and developing a deep understanding of how BAM's technology supports the investment business.
Outside of rotations, production becomes your product. The focus is on engineering solutions that make the production estate healthier, more resilient, and easier to operate - from real-time, event-driven services that detect, respond to, and resolve issues, to tooling, automation, and improvements delivered alongside our software engineering teams.
The role calls for strong coding skills, systems thinking, and operational experience to identify opportunities for improvement and build the capabilities that make them happen.
The Role in Practice
Live Production
- Own support for critical business applications across trade booking, position management, P&L, accounting, reference data, treasury, and related Operations and Finance workflows.
- Define, measure, and report SLIs, SLOs, and SLAs, using trend analysis to identify emerging reliability risks before they affect users; proactively investigate service-health anomalies, error-budget consumption, capacity signals, and recurring incident patterns to drive prompt remediation and sustained reliability improvements.
- Partner with engineering and operations teams to set up actionable alerting, reduce noise, and embed reliability targets into service design, delivery.
- Operate in a global follow-the-sun support model, providing effective regional handovers and taking part in rotational evening and weekend escalation support where needed.
- Protect critical daily business processes by ensuring the prompt and accurate completion of end-of-day workflows, including P&L delivery, accounting processes, reconciliations, and downstream data loads. Investigate exceptions and see them through to resolution.
- Troubleshoot complex production issues end to end across applications, databases, infrastructure, integrations, and business workflows. Coordinate technical and business stakeholders to restore service quickly, communicate clearly, and prevent recurrence.
- Partner closely with business teams across Operations, Finance, Treasury, Portfolio Management, and middle- and back-office functions to understand priorities, resolve issues, and improve the services they rely on.
- Support critical integrations and connectivity, including vendor data feeds, APIs, messaging platforms, FIX connectivity, managed file transfers, and secure vendor or counterparty onboarding.
- Deliver safe and well-managed change by supporting product launches, system releases, user setup, testing, implementation, and production readiness activities.
Engineering For Production
- Turn production experience into engineering solutions. Use insights from live production to identify opportunities and build capabilities that improve the health, resilience, and operability of our technology estate. Work across a broad and evolving landscape, including cloud-hosted and on-premises systems, event streaming, message buses, container orchestration, and FaaS architectures, with opportunities to work in Rust, C#, Python, and other technologies.
- Apply AI and automation to production knowledge, incident history, diagnostics, and remediation so that recurring issues are detected earlier, resolved faster, and increasingly handled without manual intervention. Build or integrate agents and workflows that assist with live production investigation and resolution.
- Design, implement, and continuously improve observability and resilience capabilities-including monitoring, alerting, dashboards, telemetry, traceability, recoverability, and operational readiness through focused sprints and project delivery to ensure the reliability and performance of critical platforms and services.
- Solve problems wherever the solution takes you. Improve existing features, automate repetitive processes, build real-time services and operational tooling, or work in unfamiliar codebases to diagnose issues and deliver permanent fixes.
What You'll Bring
- Previous experience as a Production Engineer, Site Reliability Engineer (SRE), Software Engineer, or in a similar hands-on engineering role is required.
- Experience building observability and monitoring using OpenTelemetry, Grafana observability stack (Prometheus, Loki, and Alertmanager) and ITRS Geneos.
- Experience building and maintaining software in a modern programming language. Our engineering estate includes Python, Rust, C#, and other languages. Experience in any comparable language is equally valued.
- Experience supporting financial systems or working within an investment manager or financial technology provider is desirable but not essential. A willingness and ability to learn the business domain - including trading, P&L, accounting, treasury, and operational workflows is essential.
- Experience with production controls, reconciliations, or exception-management workflows is beneficial. The ability to understand and improve complex business workflows is essential.
- Experience with SQL and relational databases, including Microsoft SQL Server, is required; familiarity with NoSQL technologies is beneficial.
- Hands-on experience with production systems involving relational databases and one or more modern infrastructure patterns, such as cloud platforms, event streaming, message buses, containerized services, or distributed applications.
- Strong troubleshooting and problem-solving skills, with the ability to investigate and support internally developed applications across application, database, infrastructure, integration, and workflow layers.
- Experience building, operating, and monitoring production services in a major cloud environment, such as AWS, Microsoft Azure, or Google Cloud Platform (GCP).
- Relevant college degree or equivalent professional experience.
- Demonstrated ability to independently design, build, test, deploy, and operate software solutions that address production or operational problems.
What We Value
- Ownership of problems through resolution and a focus on measurable business outcomes.
- Clear, direct communication and strong collaboration across teams and regions.
- A commitment to knowledge sharing and continuous improvement.
- The ability to prioritize effectively and perform well in a fast-paced trading environment.
- A culture-building mindset that helps develop people and strengthens global support practices.
- A pragmatic delivery mindset, with the ability to move quickly from a production problem to a safe, maintainable solution and deliver measurable improvements in days or weeks rather than waiting for larger platform programs.