DiligenceVault
RegTech (Regulatory & Compliance Technology)
See how your profile stacks up against this role.
We compared the job requirements to your profile to show where you're strong and where you fall short.
Job title: Senior Data Engineering Consultant -Platform Architecture & AI-Native Data Strategy
Engagement type: Contract / Consulting (3–6 months with ongoing advisory)
Location: Remote (reasonable 3-4 hrs overlap with US working hours required)
Experience: 8-10+ years in data engineering and data platform architecture
About DiligenceVault
DiligenceVault is an enterprise B2B SaaS platform that helps institutional investors, asset managers, consultants, and fund service providers digitize and automate the end-to-end due diligence lifecycle. The platform supports workflows including DDQs, RFPs, Operational Due Diligence (ODD), manager research, ESG, compliance, and investor reporting through AI-powered document processing, workflow automation, analytics, and collaboration. Today, DiligenceVault serves 100,000+ platform users, 20,000+ managers, and 250+ client teams across 150+ countries.
Our platform processes large volumes of structured and unstructured data from customer-uploaded documents, digital questionnaires, CRM systems, enterprise content repositories, regulatory filings, and platform-generated workflow data. Our current technology stack includes Azure SQL/SQL Server, Python/Celery, .NET REST APIs, Elasticsearch, Azure OpenAI, and Kestra for orchestration.
We want to build a deliberate data platform that turns this raw data into meaningful customer intelligence. We need a consultant who can help us understand the full landscape of data engineering (traditional and AI-native), assess where we are, and architect where we need to go.
What you will do
Phase 1- Educate and assess
Teach our leadership and senior architects the full spectrum of data engineering, covering traditional foundations and AI-native approaches in depth. This is not a surface-level overview - our team needs to understand concepts deeply enough to make architectural decisions. Topics span ingestion patterns (batch, streaming, CDC, adaptive connectors), transformation (ETL/ELT, dbt, Spark, LLM-assisted mapping), data modeling (dimensional, data vault, lakehouse, schema-on-read), data quality (rule-based vs. ML-driven anomaly detection, data contracts), entity resolution and data stitching (manual mapping vs. embedding-based semantic matching, knowledge graphs), orchestration (DAG engines, event-driven, self-healing pipelines), semantic layers (ontologies, contextual meaning, embedding-based search), and AI-native versioning (prompts, models, thresholds, reproducibility).
Assess our current data infrastructure end to end. Map existing data flows, identify gaps and technical debt, and produce a landscape assessment with current state, target state, gap analysis, and a prioritized roadmap.
Phase 2 - Architect and define use cases
Design the target data platform architecture across ingestion, transformation, storage, serving, and observability layers. Within this architecture, four strategic initiatives require specific attention:
PostgreSQL migration and multi-workload architecture. We are planning to move from SQL Server to PostgreSQL. The consultant will help architect a PostgreSQL environment that supports multiple workload types through the PostgreSQL extension ecosystem - pgvector for vector similarity search and embedding storage powering our AI features, analytical query patterns (columnar extensions like Citus or pg_analytics, or appropriate separation of OLAP workloads), and transactional queries for the core application. This includes guidance on connection pooling (PgBouncer/PgCat), read replica topology, partitioning strategies, and how to handle workloads that on SQL Server relied on specific features (stored procedures, Query Store, tempdb patterns) that work differently in PostgreSQL. The migration path itself - phased cutover strategy, dual-write/shadow-read validation, query translation, and performance benchmarking - is a key deliverable.
Canonical data architecture across heterogeneous sources. Data arrives from dozens of sources in different formats, schemas, and semantics- the same entity (a firm, fund, person, question) appears differently across CRM records, uploaded documents, API feeds, public filings, and form responses. The consultant will design the canonical data layer that resolves these into a unified, trustworthy representation. This covers entity resolution (how "J.P. Morgan Asset Management" in Salesforce, "JPMAM" in a DDQ, and "JPMorgan Funds" in a filing become one canonical entity), schema alignment (mapping "AUM" vs. "total_net_assets" vs. "assets_under_management" across sources), conflict resolution (when two sources disagree on a value, which wins and why), temporal alignment (different sources update at different frequencies), and the master data store that maintains these mappings with versioning and auditability. The architecture should specify where AI-native approaches (embedding-based matching, LLM-assisted semantic mapping) add genuine value vs. where traditional deterministic rules remain more reliable.
Data residency architecture for a multi-tenant, data-sharing platform. DiligenceVault is a two-sided platform where allocators (investors) and asset managers exchange due diligence data. Unlike a simple multi-tenant system where each tenant's data stays isolated, our tenants actively share data with each other- an asset manager responds to an allocator's questionnaire, and that response data is now accessible to both parties. The consultant will design a residency architecture that handles owner-anchored residency (data is stored in the region dictated by the owner's jurisdiction, even when shared with tenants in other regions), cross-region data access (a US allocator viewing responses from an EU-based manager whose data resides in the EU region), regulatory mapping (which jurisdictions require what - GDPR data localization, potential US federal requirements, APAC variations), read-path routing (queries must resolve to the correct regional store transparently), and the interaction between residency rules and features like search indexing, AI processing, and analytics aggregation.
Data governance architecture. Design the governance framework that spans the entire data platform access control (role-based and attribute-based, with tenant-level and feature-level granularity), data classification (automated PII detection, sensitivity tagging, financial data categorization), lineage and auditability (end-to-end traceability from source to served insight, including AI-model provenance), retention and lifecycle management (how long data is kept, when and how it's purged, and how regulatory holds interact with deletion requests), consent and data rights (handling data subject access requests, right-to-deletion across a shared-data platform where the same data point may be referenced by multiple tenants), quality accountability (who owns data quality at each stage, and how quality metrics are tracked and reported), and security controls (encryption at rest and in transit, key management, network isolation, and audit logging). The governance architecture should be practical for our scale - enforceable by a focused engineering team, not a 200-page policy document that nobody follows.
Define and prioritize use cases that the data platform should enable- cross-source intelligence (comprehensive fund/manager profiles by combining DDQ responses, filings, CRM data, and documents), customer behavioural insights (usage patterns, workflow analytics), automated data enrichment (linking internal data to public sources), semantic search across all documents and data, compliance signal detection (inconsistencies across a manager's submissions), and analytics/reporting pipelines serving data scientists, analysts, marketers, and client success.
Produce lasting reference materials - architecture decision records, data flow diagrams, tool evaluation guides, migration runbooks, and training decks the team uses independently after the engagement.
What we are looking for
Must have:
Strong to have:
What this role is not
This is not a staff augmentation role - the primary deliverables are knowledge transfer, architectural guidance, and strategic documents, not production code. You do not need to have used every tool mentioned - we value depth in core concepts plus the ability to quickly evaluate new tools. You do not need a research background in AI/ML, just practical understanding of how AI is changing data engineering. Financial services experience is not required but meaningfully reduces ramp-up time.
Engagement structure
Phase 1: Education, current-state assessment, landscape mapping. Deliverables: training sessions, assessment document, gap analysis, prioritized roadmap.
Phase 2: Architecture design (PostgreSQL migration plan, canonical data layer, residency architecture, governance framework), use-case specifications, tool recommendations, and reference materials.
Ongoing advisory (based on output): Periodic architecture reviews, design consultations, and progress check-ins during execution.
Time commitment - 4-5 hrs during initial phase with availability during US working hours (6-11pm IST) required. Time commitments can be reviewed later and be extended if needed.
Why Join Us?
After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.
Marcus Rivera
Chief Revenue Officer

Kodify Media Group

Diverse Lynx

Keyloop

DiligenceVault

Experlogix

DiligenceVault

DiligenceVault

DiligenceVault