- Experience
- 4+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 2 days ago
- Work mode
- Work from home
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Chipforge
Chipforge is a platform that transforms design intent into functional, validated register-transfer level (RTL) hardware, initially targeting FPGAs and aiming to expand to ASICs. This approach significantly reduces traditional development time, iteration costs, and tooling complexity. Chipforge operates as a primary strategic subsidiary of Pathkey Limited (ASX: PKY), an Australian publicly listed applied-AI company, through two entities: Chipforge Technologies Pty Ltd in Australia and Chipforge Private Limited in Singapore.
Role Overview
This position focuses on constructing the agent layer of the product—a system that interprets a design's description, generates HDL code, verifies it through rigorous tooling, and iteratively refines the output until correctness is assured. The role primarily involves working with Python and FastAPI to maintain a multi-stage context pipeline. You will take technical leadership of the agent modules, setting their strategic direction rather than performing minor tasks, collaborating closely with fellow agent engineers.
Importantly, this is a software engineering role centered on implementing robust production systems around existing AI models rather than training AI models themselves. The domain benefits from verifiable results, as designs are conclusively either successful or not upon simulation.
Key Responsibilities
- Develop new agent functionalities, including additional tool integrations, specialized agent roles, and enhancements to the request lifecycle.
- Evolve the platform’s agent patterns as the field advances, focusing on improved planning, delegation, memory management, structured tool usage, and adapting to changing provider APIs.
- Enhance context management strategies such as selection per interaction, truncation, prioritization under token limitations, and ensuring measurable quality improvements.
- Optimize cost and response latency per request without compromising output quality.
- Increase robustness of agent loops through effective error handling, retry mechanisms, fail-safes, and graceful degradation processes.
- Design and implement evaluation tools to detect regressions proactively before impacting customers.
- Implement extensive tracing and logging to facilitate debugging of non-deterministic failures.
- Scale long-duration requests supporting resumption, durable state management, and concurrent multi-instance operation.
Required Qualifications and Skills
- A minimum of four years of professional experience programming in Python, including usage of FastAPI or equivalent asynchronous frameworks. Experience working with streaming, concurrent, and long-running agentic loops is essential.
- Expertise in resilience engineering concepts such as retries, idempotency, backpressure management, graceful degradation, and reconnect-and-resume capabilities.
- Strong understanding of stateful service design capable of streaming and distributed operation across multiple instances.
- Proven production experience with agentic development focusing on tool invocation, structured outputs, streaming, retry handling, token budgeting, and managing unexpected AI model results.
- Familiarity with Server-Sent Events (SSE) or WebSocket streaming mechanisms, including handling common failure modes like partial writes, client drop-offs, and reconnections expecting session resumption.
- Ability to deliver well-designed APIs with clearly defined, stable contracts adaptable amid ongoing changes in dependent systems.
- Competency in TypeScript sufficient to engage in cross-repository changes involving streaming and tool execution.
- Experience with PostgreSQL including schema design, migration strategies, and performance optimization as data scales.
- Ability to thrive amidst architectural ambiguity, making informed design decisions without prescriptive playbooks.
- Strong written communication skills to support distributed asynchronous collaboration within a geographically dispersed team.
Desirable Skills
- Experience in agent or code generation evaluation, encompassing test harness creation, pass criteria definition, performance measurement, and regression detection.
- Hands-on knowledge of local or self-hosted inference engines (e.g., vLLM, SGLang, Ollama), along with API and system development around self-hosted AI models.
- Expertise in retrieval and ranking techniques, including Retrieval-Augmented Generation (RAG) over codebases.
- Proficiency in both TypeScript and Go to confidently make cross-layer and cross-runtime system judgments.
- Curiosity or exposure to hardware design workflows or EDA tools such as Verilog simulation, linting or synthesis, though a hardware background is not mandated.
- Experience deploying applications in air-gapped or on-premises environments.
- Background in AWS architecture and infrastructure-as-code practices, including CI/CD pipelines and production observability via logging, tracing, and alerting focused on backend or AI inference systems.
- Practical React and TypeScript skills enabling end-to-end feature delivery connecting front-end interfaces to backend services.
Working Environment
This role is remote-first with candidates welcomed from Australia or Singapore. The team is distributed across Australia, Singapore, and India, requiring ease in collaborating across time zones. While most work is performed remotely, occasional travel to collaborative team or product sessions may occur.