Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 7
ChatGPT @ 6
Communication @ 7
Data Pipelines @ 7
FastAPI
LLM @ 4
Machine Learning @ 4
Python @ 7
Salesforce @ 4
Spark @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Join Collibra's Unstructured AI team to shape how AI systems retrieve, structure, and leverage context for accurate, high-quality results at scale. Own end-to-end technical delivery of unstructured AI systems, from feature prototyping through stable production deployment across enterprise environments. Build and scale full-stack systems that ingest, process, and enrich large volumes of unstructured content, including PDFs, contracts, reports, and other document types.
The role involves collaborating with founders and cross-functional teams to understand complex business challenges, delivering solutions, and staying current with developments in machine learning and AI. This is a hybrid role based in the New York office, with at least two days per week required in the office.
Responsibilities
- Ship complex systems under ambiguity while balancing speed and precision in real-world environments.
- Write and review production-grade backend code using Python and FastAPI.
- Build and deploy document-processing systems that handle large-scale unstructured data environments.
- Integrate data from diverse enterprise sources, including SharePoint, Salesforce, and internal APIs, to provide context for AI features.
- Work with LLM-based and AI-driven enrichment capabilities such as classification, entity extraction, deduplication, and PII detection.
- Partner across engineering, product, and sales teams to ensure alignment from prototype through rollout.
- Occasionally contribute to modern frontend development.
- Develop high-performance data pipelines and context-engineering capabilities for enterprise-grade AI product features.
- Apply model evaluation best practices and search-relevance expertise.
Requirements
- Strong proficiency in Python for data processing, API development, and integrations.
- Experience delivering production-grade systems using Big Data frameworks such as Spark.
- Strong understanding of data pipelines, microservice architecture, and API design.
- Experience ingesting and processing data from third-party enterprise sources, including SharePoint/OneDrive, Salesforce, and SaaS-based knowledge bases.
- Hands-on experience with LLM-based and AI-driven enrichment models.
- Familiarity with metadata systems, data cataloging, or document AI workflows.
- Experience with search relevance.
- Strong communication skills across technical, business, engineering, product, and field teams.
- Calm, structured decision-making under tight timelines or ambiguity.
- Ability to identify risks early and course-correct effectively.
- Strong focus on data quality, precision, and governance.
- Demonstrated proficiency using AI tools such as Claude, Gemini, ChatGPT, or Copilot to solve business challenges, drive measurable outcomes, or streamline workflows.
- Bachelor's degree or equivalent related work experience.
- This position is not eligible for visa sponsorship.
Measures of Success
- Within the first month, develop a deep understanding of the product vision and unstructured data stack while shipping an initial set of end-to-end features.
- Within the third month, take ownership of technical delivery for key product areas and build robust capabilities for complex document processing and diverse data sources.
- Within the sixth month, drive ambitious enterprise-grade AI product features and architect high-performance pipelines for accurate and reliable results.
Benefits
- Competitive total rewards package.
- Bonus potential.
- Equity for eligible roles.
- Flex Fund monthly stipend.
- Pension/401k plans.
- Health coverage and time off.
- Flexible benefits designed to support employees and their loved ones.
- Equal opportunity employer with accommodations available for applicants.