Forward
AI Engineer
About this role
Forward seeks an AI Engineer to build and refine intelligent agents on top of their network digital twin platform. You'll focus on eval-driven iteration, prompt engineering, and context management to ship high-quality agentic features serving Fortune 500 companies with mission-critical infrastructure.
What you'll do
- Improve agent quality through eval-driven iteration, error analysis, and regression testing on real production trajectories
- Build new agent capabilities, tools, and interfaces with evaluation frameworks established before shipping
- Practice context engineering including retrieval optimization, prompt structuring, and token-budget management against large network models
- Engineer clear, unambiguous prompts and tool descriptions that communicate precisely with LLMs
- Define and measure success metrics for complex end-to-end agent tasks
- Collaborate with network domain experts to build and validate evaluation systems
What they're looking for
- LLM agent development and deployment
- Prompt engineering and context optimization
- Evaluation frameworks and error analysis methodologies
- Python or other production programming languages
- Software engineering fundamentals (testing, debugging, code quality)
- RAG/retrieval systems or LLM-powered workflows
- Large codebase navigation and rapid onboarding
- Technical communication and cross-functional collaboration
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Forward
Forward Networks builds a network digital twin platform designed for enterprise clients to optimize and manage their network infrastructure. The company is hiring Customer Success Engineers and other roles to provide technical expertise, drive platform adoption, and serve as primary technical contacts for customers across the country.
View all jobs at ForwardLikely interview questions
- Walk us through a shipped LLM feature you owned—how did you measure whether it worked, and what eval or quality metric drove your decisions?
- Describe a time you debugged a failing agent trajectory in production. How did you isolate the root cause and iterate on a fix?