Agent 工程

Milvus 實操

用一套一致的集合名和真實文本 Embedding,跑通寫入、檢索與刪除;程式碼假定本地 Milvus 已在 19530 端口運行。

先按 Milvus 官方 Docker Standalone 指南啓動服務,再安裝 pymilvus 與 sentence-transformers。示例使用公開文本模型並從實際輸出推導維數,不用隨機向量偽裝語義效果。
1. 連接、Embedding、Schema 與批量寫入
from pymilvus import MilvusClient, DataType
from sentence_transformers import SentenceTransformer

COLLECTION = "support_kb"
client = MilvusClient(uri="http://localhost:19530")
print(client.list_collections())  # 先確認服務可達;健康檢查常見於 9091/healthz
RESET_LAB = False  # 只有確認可丟棄舊練習數據時才改為 True
if client.has_collection(collection_name=COLLECTION):
    if not RESET_LAB:
        raise RuntimeError("support_kb 已存在;請換名稱,或確認後啓用 RESET_LAB")
    client.drop_collection(collection_name=COLLECTION)  # 刪除整個集合
encoder = SentenceTransformer("BAAI/bge-m3")
docs = [
  {"id": 1, "text": "退款通常在三個工作日內到帳", "category": "refund"},
  {"id": 2, "text": "修改密碼後需要重新登入", "category": "account"},
]
vectors = encoder.encode([d["text"] for d in docs], normalize_embeddings=True).tolist()
dim = len(vectors[0])

schema = MilvusClient.create_schema(auto_id=False, enable_dynamic_field=False)
schema.add_field(field_name="id", datatype=DataType.INT64, is_primary=True)
schema.add_field(field_name="vector", datatype=DataType.FLOAT_VECTOR, dim=dim)
schema.add_field(field_name="text", datatype=DataType.VARCHAR, max_length=1000)
schema.add_field(field_name="category", datatype=DataType.VARCHAR, max_length=64)
client.create_collection(collection_name=COLLECTION, schema=schema)

rows = [{**d, "vector": v} for d, v in zip(docs, vectors)]
client.insert(collection_name=COLLECTION, data=rows) # 數據較多時分批寫,並記錄失敗批次
2. 建索引、Load、Top-K Search 與 Filter
index = client.prepare_index_params()
index.add_index(field_name="vector", index_type="HNSW", metric_type="COSINE",
                params={"M": 16, "efConstruction": 200})
client.create_index(collection_name=COLLECTION, index_params=index)
client.load_collection(collection_name=COLLECTION)

query_vector = encoder.encode(["退的錢多久能回來?"], normalize_embeddings=True).tolist()
hits = client.search(
  collection_name=COLLECTION, data=query_vector, anns_field="vector", limit=3,
  filter='category == "refund"', output_fields=["text", "category"],
  search_params={"metric_type": "COSINE", "params": {"ef": 64}},
)
for hit in hits[0]:
    print(hit["id"], hit["distance"], hit["entity"]["text"])
3. Query 與 Delete
# Query 不做向量相似度計算
rows = client.query(collection_name=COLLECTION, filter='category == "account"',
                    output_fields=["id", "text"])
client.delete(collection_name=COLLECTION, filter="id == 2")

# ⚠️ 僅用於可丟棄的練習環境;drop 會刪除整個集合
# client.drop_collection(collection_name=COLLECTION)

寫入策略

批量 insert;使用穩定主鍵,並用 upsert 或業務去重實現冪等(insert 本身不會自動去重);保留模型名、模型版本與內容版本。更新內容時重新生成向量。

刪除策略

需要可撤銷刪除時,可先用 active=false 做邏輯刪除,並在搜索 filter 中排除;物理刪除後空間回收依賴 Compaction,不會立刻縮小文件。

檢查點 同一模型、同一維數、同一 COSINE metric、真實文本向量、先 load 後 search。先用 FLAT 做小規模正確性基綫,再比較 HNSW 的 Recall 與延遲。