Staff Engineer - Data Engineer
- Full-time
- Service Region: UCC
Company Description
We are a Digital Product Engineering company that is scaling in a big way! We build products, services, and experiences that inspire, excite, and delight. We work at scale — across all devices and digital mediums, and our people exist everywhere in the world (15000+ experts across 26 countries, to be exact). Our work culture is dynamic and non-hierarchical. We are looking for great new colleagues. That is where you come in!
Job Description
Key Responsibilities
- Design logical and physical data models for unstructured and semi-structured content (documents, case artifacts, K-Slices, extracted knowledge fragments, metadata records) originating from KM pipelines such as case mining and informal knowledge capture workflows.
- Define domain boundaries and ownership for data products — determining what constitutes a discrete, reusable data product versus a raw or intermediate asset.
- Establish metadata standards and tagging taxonomies (content type, practice/domain, provenance, confidentiality, freshness, lineage) to ensure consistent classification across knowledge sources.
- Assign and enforce security and sensitivity classifications on data products in line with firm data governance, privacy, and legal/risk requirements.
- Register, document, and maintain data products in Databricks Unity Catalog, including schemas, access grants, lineage, and catalog-level metadata.
- Partner with data engineers building Databricks pipelines to ensure ingestion, transformation, and storage patterns align to the modeled domain structure.
- Collaborate with Knowledge Products, Research Products, and Architecture/Data/Technology stakeholders to align data product design with downstream consumption needs (e.g., surfacing in Sage/Glean, AI agent retrieval).
- Support privacy and legal review processes by ensuring data products are classified and documented to enable timely sign-off.
- Establish and document repeatable modeling standards/playbooks so future data products can be onboarded consistently as the KM platform scales.
Required Qualifications
- 5+ years of experience in data modeling, data architecture, or information architecture, with meaningful exposure to unstructured or semi-structured data (not purely relational/transactional modeling).
- Direct experience working in or adjacent to Knowledge Management, content management, or enterprise search domain — understands how documents, case files, or knowledge artifacts differ from standard transactional data.
- Hands-on experience with a modern data catalog; Databricks Unity Catalog experience strongly preferred.
- Demonstrated ability to define data domains and data product boundaries in a large, multi-stakeholder organization.
- Practical knowledge of metadata management: tagging schemas, taxonomies, controlled vocabularies, or ontology design.
- Understanding of data security/sensitivity classification frameworks and how they map to access control in a lakehouse environment.
- Experience partnering with data engineering teams on ingestion and pipeline design (not required to write production pipeline code, but must speak the language).
- Strong written and verbal communication skills; able to translate technical modeling decisions into business-readable rationale for KM stakeholders and governance reviewers.
Preferred Qualifications
- Experience with enterprise knowledge platforms (e.g., Glean, SharePoint, ServiceNow) or AI-powered retrieval systems.
- Familiarity with Databricks Delta Lake, Delta Sharing, or Lakehouse Federation.
- Prior experience in professional services, consulting, or a similar document/case-intensive knowledge environment.
- Exposure to Legal/Risk/Privacy review processes for data classification and access approvals.
- Background in library science, information science, or applied ontology is a plus but not required. Success Metrics (First 6–12 Months)
- Domain model and metadata taxonomy defined and adopted for at least one major KM data product line (e.g., case mining outputs, informal knowledge K-Slices).
- Data products registered and discoverable in Unity Catalog with correct security classifications applied.
- Documented, repeatable modeling standard that engineering and future modelers can apply without re-litigating domain boundaries each time.
- Reduced turnaround time on privacy/legal classification reviews due to upfront, consistent metadata and tagging.
Qualifications
Must have skills: Data Modeling (Strong), Databricks
Good to have skills: BI Schema Design - General Experience
By clicking the link above or any third-party link within this posting, you are leaving this site and going to a third-party website where the third-party website's terms and privacy policy apply