コンテンツにスキップ

ドキュメント上のチャット (RAG)

方式 A: コーパスをすべてコンテキストに入れる

Section titled “方式 A: コーパスをすべてコンテキストに入れる”

社内ハンドブック、製品 FAQ、数本の PDF など中小規模のコーパスであれば、最もシンプルな方法は全文を system プロンプトに入れることです。検索もベクトルストアも埋め込みモデルも不要です。

from openai import OpenAI
client = OpenAI(base_url="https://api.aiand.com/v1", api_key="sk-...")
with open("handbook.md") as f:
handbook = f.read()
def answer(question: str) -> str:
response = client.chat.completions.create(
model="deepseek-ai/deepseek-v4-flash",
messages=[
{
"role": "system",
"content": (
"Answer based only on the handbook below. "
"If the answer isn't in it, say so.\n\n"
f"---\n{handbook}\n---"
),
},
{"role": "user", "content": question},
],
)
return response.choices[0].message.content
print(answer("リモートワークの規定は?"))

ハンドブック全体が context_window に収まるモデルを カタログ から選んでください — deepseek-ai/deepseek-v4-flash は最大のウィンドウを最も安く使えるモデルです。

方式 B: ローカル埋め込みでの検索

Section titled “方式 B: ローカル埋め込みでの検索”

コーパスが大きすぎてコンテキストに収まらない場合は、自前で埋め込みを計算し、クエリ時に上位 K 件のチャンクを取得します。下記の例ではローカルの sentence-transformers を使用しています:

Terminal window
pip install openai sentence-transformers numpy
import numpy as np
from openai import OpenAI
from sentence_transformers import SentenceTransformer
client = OpenAI(base_url="https://api.aiand.com/v1", api_key="sk-...")
embedder = SentenceTransformer("BAAI/bge-small-en-v1.5")
CHUNKS = [
"Refunds are processed within 5 business days of approval.",
"All employees are entitled to 20 days of paid annual leave.",
"Expense reports must be submitted within 30 days of the expense.",
"Remote work is supported for engineering and design roles.",
# ... 実際にはチャンクが数千件
]
embeds = embedder.encode(CHUNKS, normalize_embeddings=True)
def retrieve(query: str, k: int = 3) -> list[str]:
q = embedder.encode([query], normalize_embeddings=True)[0]
scores = embeds @ q
return [CHUNKS[i] for i in np.argsort(-scores)[:k]]
def answer(question: str) -> str:
context = "\n\n".join(retrieve(question))
response = client.chat.completions.create(
model="deepseek-ai/deepseek-v4-flash",
messages=[
{
"role": "system",
"content": (
"Answer based only on the snippets below. "
"If the answer isn't in them, say so.\n\n"
f"{context}"
),
},
{"role": "user", "content": question},
],
)
return response.choices[0].message.content
print(answer("リモートワークの規定は?"))

numpy 配列で足りなくなったら:

  • ベクトルストア: pgvector / Pinecone / Qdrant / Weaviate などに置き換え。
  • チャンキング: 1 チャンクあたり 200–500 トークン、若干のオーバーラップを付けて分割。ページ全体を埋め込まないこと。
  • リランキング: dense 検索で上位 K 件を取得した後、cross-encoder か LLM で再ランクする。精度向上に有効。
  • 引用: 回答とともにチャンク ID を引用させる。構造化出力 と組み合わせると確実です。
  • 埋め込みサービス: ローカル埋め込みが合わない場合は、外部の埋め込み API を同じスクリプトから呼び出すこともできます。