A UK SaaS CTO we spoke to last quarter had migrated his team's RAG system from PostgreSQL pgvector to Pinecone at £2400 monthly because "everyone uses Pinecone for AI". Six months later, his vector count was 800k, query volume was 12 QPS, and the actual latency benefit over pgvector was zero. He was paying £29k annually for infrastructure that added zero measurable value versus what he had before. Meanwhile, his team was learning a second data system, dealing with sync issues between his primary Postgres and Pinecone, and facing lock-in he had not intended.
The honest answer was that he had chosen based on vendor marketing rather than workload requirements. At his scale (under 1 million vectors, under 20 QPS), PostgreSQL pgvector was genuinely the right answer and the migration to Pinecone was a mistake driven by "AI companies use vector databases" folklore. He migrated back to pgvector, saved £29k annually, and shipped faster because his RAG system was now inside the same database as the rest of his application data.
That is the vector databases explained conversation across UK and US CTOs in 2026. Vector databases have become the default architecture pitch for any AI project involving embeddings. Sometimes a dedicated vector database is the right answer. Often PostgreSQL pgvector on your existing infrastructure is enough. And the choice should be driven by measured workload, not by vendor marketing.
This article is a candid comparison guide for CTOs choosing a vector database in 2026. Plain-language explanation of what vector databases are. When you need one versus when pgvector suffices. Head-to-head comparison of Pinecone, Weaviate, Qdrant, Milvus, and PostgreSQL pgvector. Real cost bands. What we learned running vector databases across TrackVid semantic search and IELTSArena RAG.
What a Vector Database Is (Plain CTO Language)
A vector is a list of numbers representing something in high-dimensional space. Modern language models turn text into vectors (embeddings) where semantically similar text ends up geometrically close. For example, the sentences "customer wants a refund" and "user requesting money back" have embeddings that are close together, even though they share almost no keywords.
A vector database stores these embeddings and finds nearest neighbours efficiently. Given a query embedding, it returns the K most similar stored embeddings in milliseconds even at scale of billions.
Three things distinguish a vector database from a traditional database.
Similarity search primitive. Traditional databases match exact values or ranges. Vector databases return "closest N by cosine similarity or Euclidean distance".
High-dimensional indexing. Traditional B-tree indexes fail for high-dimensional data (100-1500+ dimensions is common for embeddings). Vector databases use HNSW, IVF, or similar indexes designed for this.
Approximate nearest neighbour (ANN). Exact nearest neighbour at scale is too slow. Vector databases return approximate results (typically 95-99 percent recall) in exchange for millisecond query times.
Per Pinecone's technical documentation, well-configured vector databases return top-K similar results in under 50ms for datasets up to 100 million vectors, versus seconds to minutes for exact k-NN over the same data.
When You Actually Need a Dedicated Vector Database
Three conditions typically justify a dedicated vector database.
1. Vector count over 5-10 million with latency requirements. PostgreSQL pgvector handles up to 5 million vectors well with IVFFlat indexing. Beyond 10 million, query latency and index maintenance become noticeable. Dedicated vector databases handle 100 million to 10 billion vectors while maintaining sub-100ms latency.
2. Sustained query volume over 100+ QPS. PostgreSQL pgvector can handle bursts to 200+ QPS but sustained high QPS competes with your OLTP workload. Dedicated vector databases scale horizontally for query throughput.
3. Ops team capacity or SaaS budget available. Dedicated vector databases add infrastructure to manage (self-hosted) or subscription cost (managed SaaS). If your team is small and cost-sensitive, staying on pgvector avoids operational surface.
If all three apply, a dedicated vector database is probably the right choice. If any are missing, PostgreSQL pgvector often wins.
When PostgreSQL pgvector Is Enough
Four scenarios where pgvector on your existing PostgreSQL wins over a dedicated vector database.
Vector count under 5 million. pgvector with IVFFlat indexing performs comparably to Pinecone at this scale, at zero additional infrastructure cost.
Query volume under 50 QPS. Modern PostgreSQL handles this easily without competing with OLTP workload.
Team wants unified data model. Vectors, metadata, application data all in one database means no sync issues, transactional consistency, and one system to backup and monitor.
Cost sensitivity. pgvector on existing infrastructure adds effectively zero cost. Managed Pinecone or Weaviate starts at £70+ monthly and reaches £3000+ monthly at production scale.
Real 2026 Vector Database Comparison
Head-to-head on the criteria CTOs care about.
Feature | Pinecone | Weaviate | Qdrant | Milvus | pgvector |
Deployment | SaaS only | SaaS + self-host | SaaS + self-host | Self-host + Zilliz Cloud | PostgreSQL extension |
Starter cost | £70+/month | £0 self-host, £25+/month cloud | £0 self-host, £35+/month cloud | £0 self-host, £45+/month Zilliz | Free with Postgres |
Production scale cost | £500-£3000+/month | £150-£1500/month self-host, £300-£2500/month cloud | £150-£1200/month | £250-£1500/month | £0-£200/month additional |
Scale ceiling | 100B+ vectors | 100M+ vectors | 100M+ vectors | 10B+ vectors | 5-10M practical |
Query latency (p95) | 10-50ms | 15-60ms | 10-40ms | 15-50ms | 20-100ms |
Hybrid search | Yes | Yes (BM25 built-in) | Yes | Yes | Yes (via ts_vector) |
Filtering | Metadata filters | GraphQL + filters | Payload filters | Scalar filters | Standard SQL WHERE |
Ops burden | None (SaaS) | Medium | Low (single binary) | High (distributed) | Very low (PostgreSQL) |
Best for | High-scale production without ops team | Rich features, GraphQL fans | Simple ops, good defaults | Very high scale, ML-heavy teams | Under 5M vectors, unified data |
Which Vector Database to Choose in 2026
Choose PostgreSQL pgvector when. Vector count under 5 million. Query volume under 50 QPS sustained. Team already runs PostgreSQL. Cost sensitivity. Want unified data model without sync overhead.
Choose Qdrant when. Vector count 5M-100M. Team has some ops capacity. Want low ops burden of a dedicated vector database. Cost-conscious but need dedicated performance. Great defaults, single binary, simple operations.
Choose Weaviate when. You want built-in hybrid search (BM25 + vector). GraphQL is a plus for your team. You want SaaS or self-host flexibility. Rich metadata filtering matters.
Choose Pinecone when. You want fully managed SaaS with zero ops. Vector count over 10M with high QPS. Budget allows £500-£3000+/month for managed convenience. No self-host requirement.
Choose Milvus (or Zilliz Cloud) when. Vector count over 100M. Very high QPS requirements. Team has strong ML infrastructure background. Distributed vector search at extreme scale is the priority.
Per PostgreSQL pgvector documentation, production deployments handle 5+ million vectors with p95 latency under 100ms using IVFFlat indexing, which covers the vast majority of RAG workloads for small and mid-size SaaS products.
What We Learned Running Vector Databases Across Our Own Products
WhiteStone runs vector databases in production for TrackVid semantic search (merchant knowledge base), IELTSArena RAG (rubric-anchored teacher review), and FlexiVision similarity search. Three lessons transfer to any UK or US CTO choosing a vector database.
TrackVid started on pgvector, stayed on pgvector. TrackVid merchant knowledge search has 380k vectors and 8 QPS peak. We evaluated migrating to Pinecone in 2024 and concluded there was no measurable latency benefit at this scale. pgvector p95 latency is 45ms, which is well within acceptable range. Migration was rejected. Savings: £2400+ monthly infrastructure cost.
IELTSArena migrated from pgvector to Qdrant at 6M vectors. IELTSArena RAG vector count grew from 800k to 6M over 18 months as we ingested more IELTS band descriptors, examiner notes, and past student essays. pgvector p95 latency crept from 60ms to 210ms. Migration to Qdrant self-hosted on a £280/month dedicated instance brought p95 back to 25ms. Migration paid back within 4 months in improved teacher review experience.
We never chose Pinecone for anything. Not because it is bad. Because for our scale points (under 10M vectors typically), Qdrant self-hosted or pgvector was more cost-effective. For clients with genuinely large vector workloads (100M+), Pinecone SaaS is genuinely the right answer. We just have not had a use case that hits that scale yet.
See our portfolio of shipped work for other AI-in-production case studies. For a scoped vector database or RAG conversation, book an AI architecture call with WhiteStone.
Common Failure Modes
Choosing based on vendor marketing rather than workload. CTO migrates from pgvector to Pinecone at 800k vectors and 12 QPS because "everyone uses Pinecone for AI". Zero latency benefit, £29k annual cost. Fix: measure your actual vector count and QPS before choosing.
Underestimating pgvector at low-to-mid scale. Team assumes any AI product needs a dedicated vector database. Adds unnecessary infrastructure. pgvector would have handled the workload for years. Fix: pgvector is genuinely the right answer for under 5M vectors and under 50 QPS.
Underestimating operational complexity of Milvus. Team chooses Milvus at 3M vectors "for future scale". Distributed system operational burden dominates engineering time. Simpler choice would have been Qdrant or pgvector. Fix: match ops complexity to actual scale needs, not projected 3-year scale.
Ignoring hybrid search requirements. Team chooses pure vector search. Discovers 3 months later they need keyword + vector combined (hybrid search). Retrofit painful. Fix: evaluate hybrid search requirements from day one; choose Weaviate, Qdrant, or pgvector with ts_vector if hybrid is needed.
Frequently Asked Questions
What is a vector database and why do you need one?
A vector database stores and searches high-dimensional numeric representations (embeddings) of text, images, or other data. Instead of matching exact keywords, it finds semantically similar items. You need one when your application does semantic search, RAG, recommendations, image similarity, or duplicate detection at scale where traditional keyword search or exact matching does not work.
What is the difference between a vector database and a traditional database?
Traditional databases match exact values or ranges using B-tree indexes. Vector databases return "closest N by similarity" using specialised high-dimensional indexes (HNSW, IVF). Traditional databases are optimised for ACID transactions. Vector databases are optimised for approximate nearest neighbour (ANN) search returning results in milliseconds even at billion-scale.
Which vector database should a CTO choose in 2026?
Choose PostgreSQL pgvector for under 5M vectors and under 50 QPS. Choose Qdrant for 5M-100M vectors with low ops burden. Choose Weaviate for hybrid search or GraphQL fans. Choose Pinecone for fully managed SaaS at 10M+ vectors with budget. Choose Milvus or Zilliz Cloud for 100M+ vectors with strong ML infrastructure team. Match choice to workload, not vendor marketing.
How much does a vector database cost in 2026?
PostgreSQL pgvector on existing infrastructure: free to £200/month additional. Managed Pinecone or Weaviate Cloud: £70/month starter to £3000+/month production scale. Self-hosted Qdrant, Milvus, or Weaviate open-source: £150-£1500/month hosting. Enterprise dedicated deployments: £3000-£15000+/month for hundreds of millions of vectors with high QPS.
Do you need a dedicated vector database or is PostgreSQL pgvector enough?
pgvector is enough for vector counts under 5 million, query volumes under 50 QPS sustained, teams wanting unified data models, and cost-sensitive deployments. Dedicated vector databases become necessary above 10 million vectors, above 100 QPS sustained, or when specific features (fully managed SaaS, extreme scale, distributed vector search) are required.
When should a CTO NOT use a vector database?
Skip vector databases entirely when your data volume is small enough to fit in a language model prompt (under 20-30 documents), when your search is genuinely keyword-based (exact matching, no semantic similarity requirement), when you need strong transactional consistency across vectors and application data (pgvector wins here), or when your team lacks capacity to operate additional infrastructure.
Why choose WhiteStone Infotech for vector database and RAG architecture?
We run vector databases in production for TrackVid (pgvector at 380k vectors), IELTSArena (Qdrant at 6M vectors), and FlexiVision (Qdrant similarity search). Every architecture engagement starts with honest workload measurement (we tell you when pgvector is enough versus when dedicated is worth the operational cost). Contact WhiteStone Infotech at whitestoneinfotech.com/contact.
The One Thing to Remember
Vector database choice in 2026 should be driven by measured workload (vector count, QPS, latency requirement, ops capacity) rather than by vendor marketing. PostgreSQL pgvector is the right answer for under 5M vectors and under 50 QPS, saving £2000-£3000+ monthly versus managed SaaS. Qdrant self-hosted is the sweet spot for 5M-100M vectors with low ops burden. Pinecone SaaS is worth the cost when you have 10M+ vectors, high QPS, and no ops team. Match choice to actual scale, not projected 3-year scale, and you avoid the expensive migration mistakes that dominate this category.


.webp)
