מדריך למתחילים בנושא RAG

בדף הזה מוסבר איך להשתמש ב-Vertex AI SDK כדי להריץ את RAG Engine במשימות של פלטפורמת הסוכנים של Gemini Enterprise.

אפשר גם לעקוב אחרי ההסבר באמצעות מחברת Intro to RAG Engine.

התפקידים הנדרשים

מעניקים תפקידים לחשבון המשתמש. מריצים את הפקודה הבאה לכל אחד מהתפקידים הבאים ב-IAM: roles/aiplatform.user

gcloud projects add-iam-policy-binding PROJECT_ID --member="user:USER_IDENTIFIER" --role=ROLE

מחליפים את מה שכתוב בשדות הבאים:

  • PROJECT_ID: מזהה הפרויקט.
  • USER_IDENTIFIER: המזהה של חשבון המשתמש . לדוגמה, myemail@example.com.
  • ROLE: תפקיד ה-IAM שאתם מקצים לחשבון המשתמש.

הכנת קונסולת Google Cloud

כדי להשתמש ב-RAG Engine:

  1. מתקינים את Agent Platform SDK ל-Python.

  2. מריצים את הפקודה הזו במסוף Google Cloud כדי להגדיר את הפרויקט.

    gcloud config set project {project}

  3. מריצים את הפקודה הזו כדי לאשר את הכניסה.

    gcloud auth application-default login

הרצת מנוע RAG

מעתיקים את קוד הדוגמה הזה ומדביקים אותו במסוף Google Cloud כדי להריץ את RAG Engine.

Python

במאמר התקנת Vertex AI SDK ל-Python מוסבר איך להתקין או לעדכן את Vertex AI SDK ל-Python. מידע נוסף מופיע ב מאמרי העזרה של Python API.

import agentplatform

from agentplatform import types
from google import genai
from google.genai import types as genai_types


# Create a RAG Corpus, Import Files, and Generate a response

# TODO(developer): Update and un-comment below lines
# PROJECT_ID = "your-project-id"
# MODEL_ID = "gemini-3.5-flash"
# display_name = "test_corpus"
# gcs_path = "gs://my_bucket/my_files_dir/*"
# google_drive_path ="https://drive.google.com/file/d/123"

# Initialize Agent Platform client once per session
client = agentplatform.Client(project=PROJECT_ID, location="us-east4")

# Configure embedding model, for example "text-embedding-005".
embedding_model_config = types.RagEmbeddingModelConfig(
    vertex_prediction_endpoint=types.RagEmbeddingModelConfigVertexPredictionEndpoint(
        endpoint="publishers/google/models/text-embedding-005"
    ),
)

# Create RagCorpus
rag_corpus = client.rag.create_corpus(
    rag_corpus=types.RagCorpus(
        display_name=display_name,
        rag_vector_db_config=types.RagVectorDbConfig(
            rag_embedding_model_config=embedding_model_config
        )
    )
)

# Import Files to the RagCorpus
client.rag.import_files(
    name=rag_corpus.name,
    import_config=types.ImportRagFilesConfig(
        gcs_source=genai_types.GcsSource(uris=[gcs_path]),
        rag_file_transformation_config=types.RagFileTransformationConfig(
            rag_file_chunking_config=types.RagFileChunkingConfig(
                chunk_size=512,
                chunk_overlap=100,
            )
        ), # optional
        max_embedding_requests_per_min=1000, # optional
    )
)

# Direct context retrieval
rag_retrieval_config = genai_types.RagRetrievalConfig(
    top_k=3,  # Optional
    filter=genai_types.RagRetrievalConfigFilter(
        vector_distance_threshold=0.5
    ),  # Optional
)
response = client.rag.retrieve_contexts(
    vertex_rag_store=genai_types.VertexRagStore(
        rag_resources=[
            genai_types.VertexRagStoreRagResource(
                rag_corpus=rag_corpus.name,
            )
        ],
    ),
    query=types.RagQuery(
        text="What is RAG and why it is helpful?",
        rag_retrieval_config=rag_retrieval_config,   
    )
)
print(response)

# Enhance generation
# Create a RAG retrieval tool
rag_retrieval_tool = genai_types.Tool(
    retrieval=genai_types.Retrieval(
        vertex_rag_store=genai_types.VertexRagStore(
            rag_resources=[
                genai_types.VertexRagStoreRagResource(
                    rag_corpus=rag_corpus.name,
                    # Optional: supply IDs from `rag.list_files()`.
                    # rag_file_ids=["rag-file-1", "rag-file-2", ...],
                )
            ],
            rag_retrieval_config=rag_retrieval_config,
        ),
    )
)

# Call generate_content with the tool using the GenAI SDK

# Create a GenAI SDK client
genai_client = genai.Client(enterprise=True, project=PROJECT_ID, location="us-east4")


response = genai_client.models.generate_content(
    model=MODEL_ID,
    contents="What is RAG and why it is helpful?",
    config=genai_types.GenerateContentConfig(
        tools=[rag_retrieval_tool]
    )
)
print(response.text)
# Example response:
#   RAG stands for Retrieval-Augmented Generation.
#   It's a technique used in AI to enhance the quality of responses
# ...

curl

  1. יוצרים מאגר מידע של RAG.

      export LOCATION=LOCATION
      export PROJECT_ID=PROJECT_ID
      export CORPUS_DISPLAY_NAME=CORPUS_DISPLAY_NAME
    
      // CreateRagCorpus
      // Output: CreateRagCorpusOperationMetadata
      curl -X POST \
      -H "Authorization: Bearer $(gcloud auth print-access-token)" \
      -H "Content-Type: application/json" \
      https://LOCATION-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/LOCATION/ragCorpora \
      -d '{
            "display_name" : "'"CORPUS_DISPLAY_NAME"'"
        }'
    

    מידע נוסף זמין במאמר יצירת מאגר מידע RAG לדוגמה.

  2. מייבאים קובץ RAG.

      // ImportRagFiles
      // Import a single Cloud Storage file or all files in a Cloud Storage bucket.
      // Input: LOCATION, PROJECT_ID, RAG_CORPUS_ID, GCS_URIS
      export RAG_CORPUS_ID=RAG_CORPUS_ID
      export GCS_URIS=GCS_URIS
      export CHUNK_SIZE=CHUNK_SIZE
      export CHUNK_OVERLAP=CHUNK_OVERLAP
      export EMBEDDING_MODEL_QPM_RATE=EMBEDDING_MODEL_QPM_RATE
    
      // Output: ImportRagFilesOperationMetadataNumber
      // Use ListRagFiles, or import_result_sink to get the correct rag_file_id.
      curl -X POST \
      -H "Authorization: Bearer $(gcloud auth print-access-token)" \
      -H "Content-Type: application/json" \
      https://LOCATION-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/LOCATION/ragCorpora/RAG_CORPUS_ID/ragFiles:import \
      -d '{
        "import_rag_files_config": {
          "gcs_source": {
            "uris": "GCS_URIS"
          },
          "rag_file_chunking_config": {
            "chunk_size": CHUNK_SIZE,
            "chunk_overlap": CHUNK_OVERLAP
          },
          "max_embedding_requests_per_min": EMBEDDING_MODEL_QPM_RATE
        }
      }'
    

    מידע נוסף זמין בדוגמה לייבוא קבצים של RAG.

  3. מריצים שאילתת אחזור RAG.

      export RAG_CORPUS_RESOURCE=RAG_CORPUS_RESOURCE
      export VECTOR_DISTANCE_THRESHOLD=VECTOR_DISTANCE_THRESHOLD
      export SIMILARITY_TOP_K=SIMILARITY_TOP_K
    
      {
      "vertex_rag_store": {
          "rag_resources": {
            "rag_corpus": "RAG_CORPUS_RESOURCE"
          },
          "vector_distance_threshold": VECTOR_DISTANCE_THRESHOLD
        },
        "query": {
        "text": TEXT
        "similarity_top_k": SIMILARITY_TOP_K
        }
      }
    
      curl -X POST \
          -H "Authorization: Bearer $(gcloud auth print-access-token)" \
          -H "Content-Type: application/json; charset=utf-8" \
          -d @request.json \
          "https://LOCATION-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/LOCATION:retrieveContexts"
    

    מידע נוסף זמין במאמר RAG Engine API.

  4. ליצור תוכן.

    {
    "contents": {
      "role": "USER",
      "parts": {
        "text": "INPUT_PROMPT"
      }
    },
    "tools": {
      "retrieval": {
      "disable_attribution": false,
      "vertex_rag_store": {
        "rag_resources": {
          "rag_corpus": "RAG_CORPUS_RESOURCE"
        },
        "similarity_top_k": "SIMILARITY_TOP_K",
        "vector_distance_threshold": VECTOR_DISTANCE_THRESHOLD
      }
      }
    }
    }
    
    curl -X POST \
        -H "Authorization: Bearer $(gcloud auth print-access-token)" \
        -H "Content-Type: application/json; charset=utf-8" \
        -d @request.json \
        "https://LOCATION-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/LOCATION/publishers/google/models/MODEL_ID:GENERATION_METHOD"
    

    מידע נוסף זמין במאמר RAG Engine API.

המאמרים הבאים