Logo for Civic Marketplace

Data Engineer

Role overview

Qualifications

  • Data engineering experience in a startup or scale-up
  • Strong SQL and Python skills
  • Experience with messy external sources
  • Entity resolution or record linkage experience

Responsibilities

  • Trust in the data and maintain a golden, trusted dataset
  • Build reliable ingestion pipelines for supplier and contract data
  • Create a canonical model for supplier records
  • Develop agentic data infrastructure for AI-native procurement workflows

Key facts

Hard skills

Other skills

  • Problem Solving
  • Curiosity
  • Collaboration
  • Adaptability

About the company

Civic Marketplace logo

Civic Marketplace

GovTech & Civic Tech

Civic Marketplace is on a mission to unlock local agency procurement through innovation by building a marketplace that saves time and taxpayer dollars. Our platform provides government procurement offices with the ability to make purchases with ease, simplicity, and full legal compliance. We unlock public service delivery through cutting-edge technology, bringing transparency and efficiency. Access to a network of reliable, pre-approved vendors ensures compliance and stability for every deal, and our focus on SMWB & local economic development helps boost local growth by supporting diverse suppliers. Procurement officers face outdated technology, lacking innovation and understaffed teams. These leaders are left spending too much time on compliance and box checking when they could be focusing on helping organizations make the best sourcing decisions.. Government procurement is hindered by outdated systems and inefficient processes, leading to delays in decision-making. With limited staff and resources, these departments struggle to keep up with the increasing demand for development and public services. As a result opportunities for innovation and improved service delivery are missed, and difficulties arise in maintaining compliance and transparency. Vendors of all sizes face challenges in navigating the complex, slow and outdated procurement processes which in turn has a lasting impact on local economic development and sustainability. Vendors also struggle to demonstrate their value, as the focus is often on compliance rather than the quality and impact potential of their products and services. We have co-designed an AI-enabled marketplace that matches proven, highly-rated vendors with new, local government agency contracts. Our platform aims to combine industry best practices, legal compliance, and technology to empower procurement officers. Time savings are achieved by working with pre-approved, compliant vendors.

Company details

IndustryGovTech & Civic Tech
Company size11 - 50

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Why this role exists

Every city, county, and school district buys things. Roads, software, cleaning services, IT infrastructure. The total is somewhere north of two trillion dollars a year. And almost all of it moves through procurement processes designed for a different era: slow; paper-heavy; opaque and exhausting for everyone involved.

Civic Marketplace was built to fix that. We're a modern, data-driven platform where government agencies discover, evaluate, and engage suppliers. Where businesses, especially smaller and growing ones, can actually find and win public sector work without needing a dedicated contracts team to navigate the maze.

We're past the point of proving this works. Agencies are live on the platform, real money moves through it, and we're now combining that marketplace infrastructure with AI in ways that could genuinely transform how procurement works. Not just incrementally but structurally.

None of it works without data that can be trusted. Procurement is a data problem before it is a software problem: who the suppliers actually are, what they can genuinely deliver, which contracts an agency is already entitled to buy from, what this thing cost the county next door. That information exists, scattered across thousands of agency portals, state registries and PDF attachments, in no agreed format, and nobody has assembled it properly. Whoever does gets to define how public money is spent for the next decade. That's this role.

The problem you'd own

The hard part isn't moving data from one place to another. Any engineer here will tell you the pipelines are the easy half.

The hard part is that public procurement has no shared vocabulary. The same supplier turns up as four different legal entities across three registries, and two of the spellings are wrong. Commodity codes are applied inconsistently, or not at all. A cooperative contract one agency can buy from today is invisible to the agency next door because it was published as a PDF on a portal with no API. Every source you touch is incomplete, inconsistent, or both, and most of it is public record, so you can't quietly correct it. You have to model the mess honestly.

And the bar just moved. We recently launched an MCP integration that lets agencies request quotes through Claude, GPT and Copilot. When an agent answers a procurement question, a wrong answer doesn't read like a bug, it reads like advice, and in public spending, bad advice ends up in a council meeting. That puts weight on freshness, lineage and provenance that most product data layers never carry. As we build out our agentic procurement capabilities, the reliability of that data, and how we prepare it for retrieval, becomes our most critical engineering challenge.

So this is an entity resolution and data trust problem dressed as a pipeline problem, and it's yours to solve. You'd work out what is actually wrong with the data, which is rarely what everyone assumes, then build the fix so it holds for every source we add next. Not a script per source held together by whoever wrote it.

About the role

You would be our first dedicated data engineering hire, reporting to Mikey Mo, our Head of Engineering. Data work today sits with the product engineering team and gets done alongside shipping features. It works, but it belongs to whoever last touched it, and that isn't a foundation for what comes next.

To give you a sense of the gap: awarded quotes that slip through the platform get manually caught by our commercial team and hand-entered into HubSpot instead. Supplier onboarding data follows the same pattern. So as and when the platform and HubSpot become out of sync, it makes it difficult to say with confidence which one is right.

This is a build-and-do role. You'd collaborate closely with our Head of Engineering to shape the architecture, and you'd also write the pipelines, carry the pager for them, and go and read the raw source when a number looks wrong. If you're looking for a role where you set direction in isolation and hand implementation to other people, this isn't it. If you want substantial responsibility and scope as a function lead alongside our Head of Engineering while still enjoying the craft yourself, it very much is.

You would not be starting from nothing. There's a live platform with real transactions running through it, a product engineering team who know where the bodies are buried, and access most people in this field would have to scrape for: Civic Marketplace is a member of the NIGP Business Council, with real partnerships across councils of governments. Most people doing this job spend their first year getting the data access. You'd start with the door already open.


What you'd own

Five things, and the freedom to decide how.

  1. Trust in the data. The north star. A golden, trusted dataset that powers Civic Marketplace's analytics, so when someone asks how many quotes were awarded last quarter, or how much GMV we've captured, there's one number, and it's right. You'd own the diagnosis and the fix.

  2. Ingestion. Pipelines that pull supplier, agency, solicitation and contract data out of sources never designed to be read by anything but a person. Reliable, observable, and cheap enough to add the next source without a meeting about whether it's worth it.

  3. The canonical model. One supplier, one record, across every spelling, trading name, subsidiary and registry identifier. One shared definition of a contract vehicle, a commodity, an agency, that product, sales and customer success all use and none of them argue with. This is the unglamorous part that makes everything downstream possible.

  4. The system underneath. Testing, lineage and freshness, so when the platform tells an agency something we know where it came from and when. And self-serve data for product, customer success and analytics, so nobody queues behind an engineer for a number. Everything you build should work for the next source without you in the loop. This is where AI earns its place, turning what is currently hand-checked into something repeatable.

  5. Agentic data infrastructure. As we scale our use of AI-native procurement, you will own the data layer that powers our agentic workflows. This goes beyond standard retrieval. You will design ingestion and storage strategies, including optimizing vector pipelines like Pinecone for RAG, to ensure our agents have the high-fidelity, context-aware data necessary to perform reliably in production.

You'd work closely with product engineering, customer success and our Growth Lead. Application feature work stays with the product engineering team, so you're not competing for that ground. You own the layer they build on top of.

How we use AI

We use Claude daily, across content pipelines, community response and internal knowledge, and increasingly inside the product itself. That's real, not a line in a job ad.

For this role it cuts two ways. There's how you work: source profiling, schema mapping, test generation, the documentation that otherwise never gets written. Most of that is now automatable, and we expect you to automate it. And there's what you build: the data layer that decides whether agentic procurement is trustworthy or embarrassing. The second is the harder and more interesting problem.

What we care about is not whether you can use it. Everyone says they can. It's whether you reach for it to make something repeatable. The difference between someone who writes pipelines faster and someone who builds the tooling that means the next hundred sources don't each need hand-holding is the difference we're hiring for.

If you've built data quality tooling, evaluation harnesses, retrieval pipelines or automated documentation from scratch, tell us about them. If AI has changed how you think about the work and not just how fast you produce it, we especially want to hear that.

Our stack: Snowflake, Postgres, dbt, Fivetran, HubSpot, PostHog, Metabase.

What you bring

You'll need:

  • Data engineering experience in a startup or scale-up, where you've built the platform rather than inherited one

  • Strong SQL and Python, and the judgement to know when the warehouse is the wrong place to solve a problem

  • A data model you designed that other people had to live with, including the parts you would do differently now

  • Real experience of messy external sources: half-documented APIs, flat-file drops, scraped pages, PDFs, and data you neither control nor can correct

  • Entity resolution or record linkage experience, or clear evidence you'd be good at it. This is the centre of the job, not an edge case

  • A commercial head. You can tell the difference between a data problem that's blocking revenue and one that's merely interesting, and you sequence accordingly

  • The instinct to find out why a number is wrong before proposing a fix. Sometimes it's the pipeline, sometimes the model, sometimes the source was always like that, and those need different answers

  • The instinct to build systems that scale rather than pipelines that run once

  • Genuine AI fluency, in the sense above

  • The determination to stay positive through the ups and downs of a fast-moving startup, and to find a path through problems rather than wait for conditions to be right

  • Comfort operating remotely across timezones with real autonomy and not much oversight

  • Curiosity about why public sector data is different. You don't need to have done it. You do need to want to understand it

You'll stand out if you have:

  • Govtech, public sector or civic tech background, or experience working with public records and open data

  • Familiarity with supplier and procurement data: SAM.gov and UEI, DUNS, NAICS or UNSPSC, cooperative purchasing, COG or NIGP

  • Experience with search and retrieval, whether OpenSearch, Elasticsearch or vector pipelines (like Pinecone or Milvus) feeding an LLM product.

  • Bilingual English and Spanish, especially valuable for supplier data quality and onboarding

  • Experience being the first data hire on a team that had been doing it themselves

Why you'll love this role

The problem is genuinely important. Public procurement touches every road, school and hospital. Getting it right matters in ways most B2B SaaS doesn't. And you'd be widening access for small and local businesses, which is the part of this that keeps us up at night in a good way.

The dataset doesn't exist yet. Most data engineering jobs are being the fifth person to tidy the same warehouse. This one is assembling something nobody has assembled properly: a clean, current picture of who supplies the public sector and what public money actually buys.

The scope is unusual. First dedicated data hire, reporting to Mikey, taking substantial responsibility for the function and co-designing the architecture decisions that come with it.

You'd own a real piece of it. Equity is part of the package here, and we mean it as more than a line in an offer letter. We're looking for someone who wants to build something they have a stake in, not just a job with a good salary attached.

The AI opportunity is real. We're not bolting AI onto existing workflows. We're rethinking what's possible when procurement becomes genuinely agentic, and the data layer is the part that decides whether any of it can be trusted.

People who care. Small team, direct access to founders, genuine investment in doing things well. We move fast, but we don't move sloppy.

How we work

Build bridges to help customers win. We are obsessively focused on helping both agencies and suppliers succeed. That's not a value statement. It's the job.

High velocity, high ownership. We move quickly, make decisions, and take responsibility for outcomes. There's no one to hand things off to and no one to hide behind.

In the arena. We stay close to our users. We learn from them directly. We build based on what we see, not what we assume.

Learning quotient. Rapid iteration isn't just a process. It's a mindset. We'd rather be wrong fast and right eventually than slow and cautious throughout.

And because we work in public procurement, how we do things matters as much as what we achieve. We hold ourselves to the standards our agencies are held to.

The interview process

All interviews are via video conferencing by default.

We'll ask you to walk us through something you've built and the decisions you'd make differently now. We'll also set you a short exercise on genuinely messy source data rather than an abstract algorithm puzzle, because reasoning about real data is the job. We'll talk about how you think, not just what you've done. And we'll tell you everything we know about where Civic Marketplace is going, because we want you to be choosing us as much as we're choosing you.

What we offer

  • Competitive salary and early-stage equity

  • Comprehensive medical, dental, and vision insurance

  • Flexible PTO

  • Remote-first, with real flexibility across timezones (Remote, USA; London, UK)

  • Full AI tool stack: Claude Pro, HubSpot, Make, Notion, and more

  • Regular team offsites, including international meet-ups (ask us about Reykjavik!)

  • Direct access to the founding team and a front-row seat to building something that matters

Civic Marketplace, Inc is an equal opportunity employer. We actively encourage applications from candidates of all backgrounds, identities, and experiences.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Data Engineer Related jobs

Other jobs at Civic Marketplace

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.