Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Algorithms @ 6
CUDA
Communication @ 3
Data Science @ 3
Debugging @ 3
Distributed Systems @ 3
ElasticSearch @ 3
GPU @ 2
HPC @ 3
JVM @ 3
Java @ 3
Machine Learning @ 2
Mathematics @ 6
MongoDB @ 3
NoSQL @ 3
Performance Analysis @ 3
Software Development @ 3
Solr @ 3
Vector Databases @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for Java engineering interns to work on cuVS, a suite of open source software libraries for unstructured data processing and vector search algorithms on GPUs.
cuVS relies on NVIDIA CUDA for low-level compute optimization, but exposes that high-performance GPU compute through user-friendly languages like Java.
We're expanding our vector search and database acceleration to include a Java engineering intern. The cuVS team builds next-generation building blocks for accelerating Java-based libraries like Lucene and JVector, which are used in widely popular databases like OpenSearch, Solr, MongoDB, and Elasticsearch. In this role, you will develop, benchmark, and explore novel tuned custom solutions for accelerating vector preprocessing, clustering, indexing, and search. This includes end-to-end database acceleration and scale, including introducing scalable architectural improvements, optimizing disk-based indexing and using next-generation hardware to benchmark data volumes not tractable with today's CPU-centric solutions. You’ll work closely with the cuVS team of stellar engineering, redefining what’s possible in the world of unstructured information retrieval.
Responsibilities
- Analyze, design, and implement optimized GPU algorithms for large-scale vector search, databases, and machine learning.
- Expand and improve integration of NVIDIA cuVS into relevant high-level vector search libraries and vector databases.
- Performance analysis, benchmarking, and troubleshooting of associated libraries.
Requirements
- Currently enrolled in a Masters or PhD program in Data Science, Machine Learning or Computer Science.
- Strong analytical problem-solving skills, algorithms and mathematics fundamentals.
- Excellent software development skills: programming, debugging, performance analysis, and test design, especially within the Java ecosystem and the JVM.
- Experience with NoSQL DBs: Lucene, Elasticsearch, OpenSearch, MongoDB, Solr.
- Good communication and documentation habits.
Ways to stand out from the crowd
- Experience developing distributed algorithms and running on distributed systems: HPC, Cloud, etc.
- Distributed System experience and development.
- Experience with debugging multi-language and multi-hardware systems.
- Experience with Vector Databases: Milvus, Pinecone, LanceDB.
- Familiarity with Nearest Neighbor Algorithms like graph-based and inverted file indexes.
- Some familiarity with machine learning concepts like clustering and dimensionality reduction.
- Some familiarity with GPU programming knowledge (GPU programming knowledge is a plus; if you don’t have it, they’re happy to teach you).