- Experience
- Any
- Salary
- CAD 150,000 – CAD 250,000 / year
- Openings
- 1
- Posted
- 1 day ago
- Work mode
- Work from home
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Paires
Paires connects founders seeking capital with a global network of investors through a two-sided platform that facilitates warm outreach resulting in meetings. The company operates with a small, senior team, is profitable, self-funded, and continuously ships new features.
Role Overview
As the inaugural Data Engineer, you will take ownership of the core database powering Paires' agents and outreach systems. This database is a comprehensive knowledge graph capturing every company, investor, funding round, relevant news, and associated communications such as emails and call transcripts. Your role involves designing, scaling, maintaining data quality, and ensuring it remains the definitive source of truth for the platform. This is not a standard analytics or reporting warehouse but a live memory system tailored for precise, moment-specific fact retrieval by agents.
Key Responsibilities
- Manage and optimize the primary database built on Postgres and Supabase with hybrid search capabilities, focusing on schema design, modeling, scaling, and performance improvements. Consolidate vector data storage using pgvector rather than third-party vector databases.
- Ensure comprehensive data quality through validation processes for vendor and third-party data, deduplication, entity resolution, provenance tracking, and continuous monitoring.
- Maintain the communications layer by storing, linking, and enabling searchability of raw emails and call transcripts connected to relevant people and companies.
- Engineer ingestion and enrichment pipelines for funding rounds, market news, and extensive contact and company research, prioritizing cost-efficiency and data freshness.
- Develop and maintain the knowledge graph representing entities such as companies, investors, and funding rounds, including detailed provenance on each fact.
- Build and uphold a unified, clean data layer that supports all campaigns, agents, and product features across the platform.
Ideal Candidate Profile
- Experience managing a database comprising companies, people, deals, or their communications that serve as a CRM or intelligence source used by live products or sales teams.
- Proficiency in SQL and Python, with demonstrated experience creating data pipelines involving ingestion, transformation, deduplication, and enrichment.
- A track record of identifying and mitigating bad data issues before they impact business outcomes.
- Strong conceptual skills in schema design and contract definition, planning for scalable query needs over time.
- Expertise modeling large-scale entities and their relationships—companies to investors to funding rounds to individuals—while maintaining efficient queryability as data sources grow.
- Ability to rapidly adopt AI tools and take full ownership of related deliverables.
- Relevant hands-on experience such as managing CRM data, enrichment pipelines, deduplication, and interfacing with AI-driven processes is valuable, regardless of formal title.
- Bonus skills include familiarity with pgvector and embeddings, constructing knowledge graphs within relational databases, scalable ingestion of funding rounds or news, entity resolution at scale, and managing raw communications storage solutions.
Work Environment and Benefits
- Fully remote, asynchronous work setup with some overlap aligning with US Eastern Time for collaboration.
- Meetings are efficiently scheduled twice weekly, allowing for extensive deep work time.
- Access to top-tier AI tooling, fully paid, including Claude Code, Cursor, and advanced models.
- Collaboration with Go-To-Market leadership and founding engineers, influencing all layers of the platform.
Compensation
Salary range between CA$150,000 and CA$250,000 per year.