Nybble™
SnackStackHack
Vector Databases, Inside and Out

Lesson 14 of 18

Lessons

  1. 1. The similarity search problem
  2. 2. Distance metrics and vector spaces
  3. 3. Embedding models — choosing and evaluating
  4. 4. How embedding models are trained
  5. 5. HNSW: the production default
  6. 6. IVF: cluster-based search
  7. 7. Quantization: compressing vectors for scale
  8. 8. DiskANN and beyond: billion-scale search
  9. 9. Architecture of a vector database
  10. 10. Metadata filtering and hybrid search
  11. 11. Multi-vector search and late interaction
  12. 12. pgvector: the Postgres-native vector store
  13. 13. Choosing a vector database
  14. 14. Ingestion pipelines and corpus management
  15. 15. Scaling and multi-tenancy
  16. 16. Security and compliance
  17. 17. Operations and observability
  18. 18. Building a production vector search service
  19. Key takeaways
  20. How to get certified
  21. Your certificate

Ingestion pipelines and corpus management

Free account

Read this lesson

The opening lessons of every course are free — this one needs an account. Sign in and the full course opens.

  • Every lesson, start to finish
  • Progress saved across devices
  • Bits per lesson, plus a bonus for finishing

Free · your email is used for progress only.

Nybble™ — built for the people building AI.

AboutTermsPrivacyContact
SnackStackHack
Message Nybble