C

Software Engineer, AI Agent Platform

Chipforge.ai

Remote · Full Time

Be the first to apply

Experience
4+ yrs
Salary
Openings
1
Posted
1 day ago
Work mode
Work from home
Resume
Required to apply

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About Chipforge

Chipforge offers a platform transitioning from design intent to validated register-transfer level (RTL) hardware, initially focusing on FPGAs and expanding towards ASICs. Their approach reduces traditional cycle times, iteration expenses, and complexities associated with conventional toolchains. As the primary strategic subsidiary of Pathkey Limited (ASX: PKY), an Australian-listed applied-AI group, Chipforge operates through entities in Australia and Singapore.

Role Overview

This position involves developing the agent layer within Chipforge's product: a system that interprets design specifications, generates HDL code, validates outputs using verification tools, and iteratively refines results until verification succeeds. The technology stack centers mainly on Python and FastAPI within a multi-stage contextual pipeline. The role demands technical leadership over agent components rather than peripheral task handling and close cooperation with other agent engineers.

Key Responsibilities

  • Develop new agent functionalities including integrating additional tools, specialized agent roles, and novel stages in the request lifecycle.
  • Enhance the platform's agent patterns to incorporate advancements such as improved planning, delegation, memory, structured tooling, and evolving provider APIs.
  • Refine context management strategies focusing on content inclusion per interaction, token limit-based truncation and prioritization, aimed at measurable quality enhancements.
  • Balance cost and latency optimization per request while preserving output quality.
  • Strengthen agentic loops by implementing robust error handling, retry mechanisms, fallbacks, and graceful degradation.
  • Create evaluation frameworks that detect regressions proactively before impacting customers.
  • Implement detailed tracing and logging to facilitate debugging of non-deterministic failures.
  • Support scaling of long-duration requests with features like resume capabilities, durable state persistence, and multi-instance support.

Required Qualifications and Skills

  • A minimum of 4 years' experience in production-grade Python development, especially using FastAPI or comparable asynchronous frameworks. Experience with streaming, concurrent, and long-running agentic loops is essential.
  • Expertise in resilience engineering including techniques for retries, idempotency, backpressure management, graceful degradation, and reconnect-resume mechanisms.
  • Proven skills in designing stateful, streaming services capable of operating across multiple instances.
  • Hands-on experience with agentic development concepts such as tool invocation, structured output formatting, streaming data, retry strategies, token allocation, and comprehensive error management for unexpected model behaviors.
  • Proficiency in SSE or WebSocket streaming protocols and handling associated failure modes including partial data transmission, dropped connections, and reconnect-resume expectations.
  • Strong API design capabilities with an emphasis on establishing and maintaining stable contracts while enabling simultaneous changes across dependent codebases.
  • Working knowledge of TypeScript sufficient to contribute to streaming and tool-execution features in a cross-repository environment.
  • Experience with PostgreSQL focusing on schema design, migration processes, and optimization of query performance as dataset sizes grow.
  • Ability to operate effectively amidst architectural ambiguity, making informed design decisions without established guidelines.
  • Excellent written communication skills to support asynchronous collaboration within a distributed team.

Preferred but Not Mandatory

  • Exposure to agent evaluation practices involving harness construction, setting pass criteria, performance measurement, and regression detection.
  • Familiarity with local and self-hosted AI inference frameworks such as vLLM, SGLang, or Ollama, including building APIs and systems to support these models.
  • Experience with retrieval and ranking methods, notably retrieval-augmented generation (RAG) applied to codebases.
  • Competence in TypeScript and Go for conducting system-level assessments across multiple layers and runtime environments.
  • Interest or knowledge in hardware design workflows or electronic design automation (EDA) tools including Verilog simulation, linting, or synthesis.
  • Experience developing solutions for air-gapped or on-premise deployment settings.
  • Expertise with AWS architecture and infrastructure-as-code, encompassing CI/CD pipelines and production-grade observability such as logging, tracing, and alerting for backend or AI inference systems.
  • Practical capabilities in React and TypeScript for integrating front-end interfaces or independently delivering complete features.

Work Environment

This role is remote-first and open to candidates residing in Australia or Singapore. The team operates across multiple regions including Australia, Singapore, and India, requiring comfort with asynchronous and timezone-spanning collaboration. Occasionally, travel may be necessary for team or product gatherings, but the primary work setup is remote.

Tools & software

PostgreSQL required

How they work

Communication Teamwork & Collaboration Independence Resilience
🤖
Online · instant AI help
Broxer