AI Unit Economics, FinOps & Infrastructure Cost ModelingPlaybook3 min readUpdated September 2026

Pgvector, Pinecone or Qdrant: Comparing Real Vector Database Costs

Vector database pricing pages are hard to compare because each vendor charges for a different combination of things: storage, queries, compute, or all three bundled into a per-pod rate. The only reliable comparison is one you run yourself, against your own data shape and query pattern.

Here's what actually drives the cost in each of the three most common choices, and how to set up a fair test between them.

What actually drives vector database cost

Regardless of vendor, the cost drivers are the same handful of variables: number of vectors stored, dimensionality of each vector, queries per second at peak, and whether you need metadata filtering alongside the similarity search. A million vectors at 384 dimensions is a very different bill than a million vectors at 1536 dimensions, and query volume matters as much as storage for anything billed on compute.

Pgvector on infrastructure you already run

Pgvector adds vector search to a Postgres instance you're likely already paying for, which means no new vendor bill, but the cost shows up instead as the compute you allocate to that Postgres instance and the engineering time to tune indexes as your data grows. It tends to be the cheapest option at moderate scale and the one that requires the most of your own database expertise to keep fast as you scale up.

Pinecone and other managed services

A managed vector database removes the operational burden entirely, at the cost of a dedicated bill, usually priced by pod size or by a combination of stored vectors and query volume. This is the right tradeoff when your team's time is better spent on the product than on database tuning, or when query volume is high enough that self-managing would need dedicated infrastructure expertise you don't have in house.

Qdrant and the self-hosted middle ground

Self-hosted Qdrant sits between the two: purpose built for vector search like a managed service, but running on infrastructure you provision and operate yourself, which trades a per-vendor bill for cloud compute cost plus your own operational time. It's a reasonable landing spot for teams that want vector search performance without a managed vendor bill, and are comfortable running one more service themselves.

Running your own three way cost test

Take a representative slice of your actual data, load it into each option, and run your actual query pattern against all three, not a generic benchmark. Measure the fully loaded cost for each: infrastructure spend plus a reasonable estimate of engineering time to operate it monthly. The option that wins on a vendor's marketing page is not always the one that wins against your specific data shape and team's operational capacity.

Give the test at least a few weeks of realistic query volume, not a single afternoon, since index build time and query latency both change as data volume grows past what a quick test usually covers. A cost comparison run against a small sample can understate how a managed service's per-query pricing or a self-hosted option's compute needs actually scale once you're near production volume.

Follow these steps to run a fair three way test:

  1. Take a representative slice of your real data, stored at the dimensionality your embedding model actually produces.
  2. Load that slice into pgvector, Pinecone and Qdrant so each option holds the same data.
  3. Run your actual query pattern against all three, including any metadata filters you use in production, rather than a generic benchmark.
  4. Record infrastructure spend for each option over the same period.
  5. Add a reasonable monthly estimate of engineering time to operate each one, then compare the fully loaded totals.

Deciding when to revisit the choice

Vector database pricing changes often, and a comparison you ran a year ago may no longer hold. Revisit the decision whenever your stored vector count roughly doubles, whenever a vendor changes its pricing structure, or whenever your query volume shifts meaningfully, rather than assuming the original choice is permanent. Treat it the way you'd treat any infrastructure vendor decision: right for a given scale, not right forever.

Write down the assumptions behind whichever choice you make, specifically the vector count and query volume you tested at, so the next person reviewing the decision knows whether current usage still matches the conditions the original test was run under.

The mistake that skews most vendor comparisons

Most public benchmarks test a simple nearest neighbor lookup with no metadata filter attached, and plenty of real workloads filter by tenant, date range or category on every query. Adding a filter changes which index structures perform well and can shift the cost comparison entirely, since a database that's fast on a bare similarity search can slow down considerably once it also has to satisfy a filter on every request.

If your product filters on most queries, which is common for anything multi-tenant, run your cost test with that filter included from the start rather than adding it after you've already picked a favorite based on the unfiltered number. The unfiltered result is the one vendors tend to publish, and it's the one least likely to match how your application actually queries the database.

The same logic applies to write pattern, not just reads. A workload that inserts vectors in a steady trickle behaves differently, cost-wise, than one that bulk loads a large batch overnight and then reads heavily during the day. Test whichever pattern actually matches your product, since a vendor tuned for one pattern can look artificially cheap or expensive under the other.

Executive Capability Standard

What Good Looks Like

Good looks like a cost comparison run against your own data and query pattern, not a vendor's published benchmark.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Understand what unit each vendor actually bills on: stored vectors, queries, pod size, or compute time, before comparing sticker prices.
2. Do Manually:Load a representative sample of your real data into each candidate and run your own query pattern against all three.
3. Delegate:Have a backend engineer own the load test and cost comparison, with finance reviewing the fully loaded totals before a decision.
4. Automate:Script the load and query test so it can be rerun automatically whenever a vendor changes pricing or your data volume grows materially.
5. Buy:Choose a managed vector database outright once your team's engineering time is worth more spent elsewhere than on database operations.

How to Get Started

Frequently Asked Questions

Is pgvector always the cheapest option?

At moderate scale and if you already run Postgres, usually yes on direct dollar cost. That advantage narrows as vector count and query volume grow, since you'll eventually need database expertise or dedicated compute to keep it fast, which has its own real cost even without a separate vendor invoice.

How do we estimate the engineering time cost of self-hosting a vector database?

Track the hours actually spent on index tuning, scaling, and incident response over a real month, then apply a fully loaded hourly cost. Teams consistently underestimate this until they've run it for at least one full month past the initial setup.

Does dimensionality really change the cost that much?

Yes. Storage and compute both scale with dimensionality, so a model that produces 1536 dimension embeddings costs meaningfully more to store and search than one producing 384 dimensions, independent of which vector database you choose. It's worth checking whether a smaller embedding model meets your accuracy needs before committing to infrastructure sized for a larger one.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides