ZipNom Logo

ZipNom

Data Engineer

Posted 10 Hours Ago
Be an Early Applicant
Remote
Hiring Remotely in IND
Mid level
Remote
Hiring Remotely in IND
Mid level
Build and operate a production data pipeline ingesting Indian MCA registration data. Responsibilities include raw-first S3 storage, resilient scraping and parsing, snapshot and diff processing, granular change events, scheduled batch jobs, alert digests, monitoring, anomaly detection, reference-table maintenance, and historical backfills. The role requires strong Python, PostgreSQL, testing, operational reliability, and production experience with uncooperative data sources.
The summary above was generated by AI

This is a remote position.

About Verizol

Verizol is ZipNom's company intelligence and verification platform: 45+ APIs (KYC, KYB, bank, face, OCR, CKYC, eSign) built on a daily MCA data pipeline, sold through a single API key and prepaid credit wallet. Customers are CA firms, fintechs, NBFCs, and marketplaces. The engineering blueprint is written, the 8-week build plan is set, and the team is small enough that your name is on what ships.

The Role

You own the asset the entire company is built on. Every Verizol product — free search, daily alerts, KYB reports, watchlists — reads from one normalized, indexed database of Indian company registrations, and that database exists because your pipeline ingests, parses, and diffs MCA data every single day. This is the least substitutable seat on the team: when a government data source silently changes format at 2 AM, you're the one whose alarms catch it and whose parsers get fixed before customers notice.

This is not clean-API ETL. It's real-world data engineering against uncooperative sources — scraping, format archaeology, checksum validation, and building a system honest enough to know when its own data is stale.




Requirements

What You Will Do

  • Build source ingestors for MCA data (new incorporations, master data, index of charges, DIN registry, filing metadata) with raw-first storage to S3 — every byte stored before parsing, so parser bugs are always replayable
  • Write parsers as pure, fixture-tested functions: raw bytes in, typed rows out, with a fixture for every format variant ever observed in the wild
  • Build and own the snapshot + diff engine: detect exactly what changed for every company daily, and emit granular change events (director resigned, charge created, status changed) — the event stream that powers Alerts and Watchlists
  • Own the daily production schedule: ingest 02:00 → diff 04:00 → digests 06:30 → alert emails at 08:00 IST, with stage-level retries and a hard rule that stale data never silently ships
  • Build the alerts digest job: filter matching across all subscriber streams, email rendering, SES delivery — for CA subscribers, this email is the product
  • Build the monitoring that keeps you sane: row-count anomaly alarms, parser error rates, per-source freshness gauges
  • Maintain the reference tables (state/ROC mappings, NIC codes) and run the historical backfill (24+ months of incorporations)

Required Skills

  • 3–6 years of data engineering or backend work with production scraping/parsing experience against uncooperative sources — this is the non-negotiable; clean-API ETL alone won't prepare you for MCA
  • Strong Python: Celery (or equivalent task queues), lxml/BeautifulSoup, pandas, pdfplumber or similar document parsing
  • Solid PostgreSQL: bulk upserts, COPY, partitioning, and index design for time-series feed queries
  • Testing instincts for data: fixtures, golden files, replay tests — you can prove a parser change didn't corrupt yesterday
  • Operational maturity: your jobs run while everyone sleeps; you build the alarms first and take the 2 AM page seriously
  • Comfort with proxies, rate budgets, and being a polite, resilient client of fragile infrastructure

Preferred

  • Prior experience with Indian government/registry data: MCA21, GSTN, EPFO, court records, or similar
  • Familiarity with CIN/DIN structures, ROC organization, or corporate filings
  • AWS: S3 lifecycle policies, SQS, CloudWatch metrics/alarms
  • Email deliverability basics (SES, DKIM) — you'll co-own the digest send


Similar Jobs

6 Days Ago
In-Office or Remote
Expert/Leader
Expert/Leader
Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Lead data engineering architecture and execution for large-scale, low-latency analytics. Own technical direction, ensure operational data quality, design streaming and batch pipelines, mentor engineers, coordinate cross-functional teams, and deliver scalable solutions using Spark, Airflow, and AWS or equivalent data platforms.
Top Skills: AirflowAthenaColumn StoresDatabricksDatabricks ApisEmrFlinkHiveKafkaKappa ArchitectureLambda ArchitectureMaster Data Management (Mdm)Microservices ArchitectureRedshiftSparkSQLStreaming Pipelines
9 Hours Ago
Remote
IND
Mid level
Mid level
HR Tech • Information Technology • Professional Services
Design, implement, and optimize graph data solutions for a supplier chain project. Responsibilities include graph modeling with LPG and RDF, Neo4j database development and query optimization, GCP data processing using BigQuery, GCS, and Dataproc, and integrating graph data with visualization tools. The role also involves schema and data design, SQL-based relational database work, performance tuning, and potentially ETL/ELT, streaming, API, and distributed systems development.
Top Skills: Api DevelopmentBigQueryCypherDataprocDistributed SystemsEtl/EltGcp Data FusionGoogle Cloud Platform (Gcp)Google Cloud Storage (Gcs)KafkaKinesisLabeled Property Graphs (Lpg)Neo DashNeo4JNeo4J AuraPower BIPysparkResource Description Framework (Rdf)SparksqlSQLTableau
Yesterday
Remote
IND
Senior level
Senior level
HR Tech • Information Technology • Professional Services
Develop Azure data pipelines and Databricks Lakehouse solutions using Azure Data Factory, Azure SQL, SQL, and PySpark. Convert Informatica PowerCenter and Oracle PL/SQL code, curate and transform structured and unstructured data, implement data models and Delta Lake tables, and support SQL Server cloud migration. Use Azure services including ADLS and ADX, automate delivery through DevOps and CI/CD, estimate project effort, and lead complex development in an agile environment.
Top Skills: AdlsAzure Data ExplorerAzure Data FactoryAzure Data Lake StorageAzure DatabricksAzure SqlCi/CdDatabricks LakehouseDelta LakeDelta TablesDevOpsInformatica PowercenterKiteworksOracle Pl/SqlPysparkRdbmsSQLSQL Server

What you need to know about the Delhi Tech Scene

Delhi, India's capital city, is a place where tradition and progress co-exist. While Old Delhi is known for its rich history and bustling markets, New Delhi is defined by its modern architecture. It's clear the region places a strong emphasis on preserving its cultural heritage while embracing technological advancements, particularly in artificial intelligence, which plays a central role in shaping the city's tech landscape, fueled by investments in research and development.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account