DiligenceVault

Senior Data Engineering Consultant -Platform Architecture & AI-Native Data Strategy

Remote · Posted 4d ago

salary not listedseniorcontractremote
Azure SQLSQL ServerPythonCelery.NETElasticsearchAzure OpenAIKestradbtSparkPostgreSQL

Job Description

Job title: Senior Data Engineering Consultant -Platform Architecture & AI-Native Data Strategy Engagement type: Contract / Consulting (3–6 months with ongoing advisory) Location: Remote (reasonable 3-4 hrs overlap with US working hours required) Experience: 8-10+ years in data engineering and data platform architecture About DiligenceVault DiligenceVault is an enterprise B2B SaaS platform that helps institutional investors, asset managers, consultants, and fund service providers digitize and automate the end-to-end due diligence lifecycle. The platform supports workflows including DDQs, RFPs, Operational Due Diligence (ODD), manager research, ESG, compliance, and investor reporting through AI-powered document processing, workflow automation, analytics, and collaboration. Today, DiligenceVault serves 100,000+ platform users, 20,000+ managers, and 250+ client teams across 150+ countries. Our platform processes large volumes of structured and unstructured data from customer-uploaded documents, digital questionnaires, CRM systems, enterprise content repositories, regulatory filings, and platform-generated workflow data. Our current technology stack includes Azure SQL/SQL Server, Python/Celery, .NET REST APIs, Elasticsearch, Azure OpenAI, and Kestra for orchestration. We want to build a deliberate data platform that turns this raw data into meaningful customer intelligence. We need a consultant who can help us understand the full landscape of data engineering (traditional and AI-native), assess where we are, and architect where we need to go. What you will do Phase 1- Educate and assess Teach our leadership and senior architects the full spectrum of data engineering, covering traditional foundations and AI-native approaches in depth. This is not a surface-level overview - our team needs to understand concepts deeply enough to make architectural decisions. Topics span ingestion patterns (batch, streaming, CDC, adaptive connectors), transformation (ETL/ELT, dbt, Spark, LLM-assisted mapping), data modeling (dimensional, data vault, lakehouse, schema-on-read), data quality (rule-based vs. ML-driven anomaly detection, data contracts), entity resolution and data stitching (manual mapping vs. embedding-based semantic matching, knowledge graphs), orchestration (DAG engines, event-driven, self-healing pipelines), semantic layers (ontologies, contextual meaning, embedding-based search), and AI-native versioning (prompts, models, thresholds, reproducibility). Assess our current data infrastructure end to end. Map existing data flows, identify gaps and technical debt, and produce a landscape assessment with current state, target state, gap analysis, and a prioritized roadmap. Phase 2 - Architect and define use cases Design the target data platform architecture across ingestion, transformation, storage, serving, and observability layers. Within this architecture, four strategic initiatives require specific attention: PostgreSQL migration and multi-workload architecture. We are planning to move from SQL Server to PostgreSQL. The consultant will help architect a PostgreSQL environment that supports multiple workload types through the PostgreSQL extension ecosystem - pgvector for vector similarity search and embedding storage powering our AI features, analytical query patterns (columnar extensions like Citus or pg_analytics, or appropriate separation of OLAP workloads), and transactional queries for the core application. This includes guidance on connection pooling (PgBouncer/PgCat), read replica topology, partitioning strategies, and how to handle workloads that on SQL Server relied on specific features (stored procedures, Query Store, tempdb patterns) that work differently in PostgreSQL. The migration path itself - phased cutover strategy, dual-write/shadow-read validation, query translation, and performance benchmarking - is a key deliverable. Canonical data architecture across heterogeneous sources. Data arrives from dozens of sources in different formats, schemas, and semantics- the same entity (a firm, fund, person, question) appears differently across CRM records, uploaded documents, API feeds, public filings, and form responses. The consultant will design the canonical data layer that resolves these into a unified, trustworthy representation. This covers entity resolution (how "J.P. Morgan Asset Management" in Salesforce, "JPMAM" in a DDQ, and "JPMorgan Funds" in a filing become one canonical entity), schema alignment (mapping "AUM" vs. "total_net_assets" vs. "assets_under_management" across sources), conflict resolution (when two sources disagree on a value, which wins and why), temporal alignment (different sources update at different frequencies), and the master data store that maintains these mappings with versioning and auditability. The architecture should specify where AI-native approaches (embedding-based matching, LLM-assisted semantic mapping) add genuine value vs. where traditional deterministic rul