זיהוי כוונות באמצעות פלט אודיו

לפעמים אפליקציות צריכות בוט כדי להשיב למשתמש הקצה. Dialogflow לפעמים אפליקציות צריכות בוט כדי להשיב למשתמש הקצה. ‫Dialogflow מאפשר לכם להשתמש ב-Cloud Text-to-Speech שמבוסס על DeepMind WaveNet כדי ליצור תשובות קוליות מהנציג שלכם. ההמרה הזו של תגובות טקסט ל-Intent לאודיו נקראת פלט אודיו, סינתזת דיבור,

אתם יכולים להשתמש באודיו גם לקלט וגם לפלט כשאתם מזהים כוונה. זה קורה בדרך כלל כשמפתחים אפליקציות שמתקשרות עם משתמשים באמצעות ממשק אודיו בלבד.

רשימת השפות הנתמכות מופיעה בעמודה TTS בדף שפות.

לפני שמתחילים

התכונה הזו רלוונטית רק כשמשתמשים ב-API עבור אינטראקציות עם משתמשי קצה. אם אתם משתמשים בשילוב, אתם יכולים לדלג על המדריך הזה.

לפני שמתחילים, צריך לבצע את השלבים הבאים:

יצירת סוכן

  1. עוברים אל מסוף Dialogflow ES.
  2. נכנסים למסוף אם מוצגת בקשה לכך. מידע נוסף זמין בסקירה הכללית על מסוף Dialogflow.
  3. בתפריט שבסרגל הצד, מרחיבים את האפשרות סוכני טעינה.
  4. לוחצים על יצירת סוכן חדש.
  5. מזינים את שם הסוכן, שפת ברירת המחדל ואזור הזמן שמוגדר כברירת מחדל.
  6. מזינים פרויקט קיים. כדי לאפשר למסוף Dialogflow ליצור פרויקט, בוחרים באפשרות Create a new Google project (יצירת פרויקט חדש ב-Google).
  7. לוחצים על יצירה.

זיהוי כוונות

כדי לזהות כוונות, מבצעים קריאה ל-method‏ detectIntent בסוג Sessions.

REST

1. הכנת תוכן אודיו

מורידים את קובץ האודיו לדוגמה book-a-room.wav, שבו נאמר "book a room". קובץ האודיו צריך להיות בקידוד base64 בדוגמה הזו, כדי שאפשר יהיה לספק אותו בבקשת ה-JSON שבהמשך.

דוגמה ל-Linux:

wget https://cloud.google.com/dialogflow/es/docs/data/book-a-room.wav
base64 -w 0 book-a-room.wav > book-a-room.b64

דוגמאות לפלטפורמות אחרות מופיעות במאמר הטמעת אודיו בקידוד Base64 במסמכי התיעוד של Cloud Speech API.

2. שליחת בקשה לזיהוי כוונת המשתמש

מבצעים קריאה ל-detectIntent בסוג Sessions ומציינים אודיו בקידוד base64.

לפני שמשתמשים בנתוני הבקשה, צריך להחליף את הנתונים הבאים:

  • ‫PROJECT_ID: מזהה הפרויקט ב-Google Cloud
  • ‫SESSION_ID: מזהה סשן
  • ‫BASE64_AUDIO: התוכן בקידוד base64 מקובץ הפלט שלמעלה

ה-method של ה-HTTP וכתובת ה-URL:

POST https://dialogflow.googleapis.com/v2/projects/PROJECT_ID/agent/sessions/SESSION_ID:detectIntent

תוכן בקשת JSON:

{
  "queryInput": {
    "audioConfig": {
      "languageCode": "en-US"
    }
  },
  "outputAudioConfig" : {
    "audioEncoding": "OUTPUT_AUDIO_ENCODING_LINEAR_16"
  },
  "inputAudio": "BASE64_AUDIO"
}

כדי לשלוח את הבקשה צריך להרחיב אחת מהאפשרויות הבאות:

אתם אמורים לקבל תגובת JSON שדומה לזו:

{
  "responseId": "b7405848-2a3a-4e26-b9c6-c4cf9c9a22ee",
  "queryResult": {
    "queryText": "book a room",
    "speechRecognitionConfidence": 0.8616504,
    "action": "room.reservation",
    "parameters": {
      "time": "",
      "date": "",
      "duration": "",
      "guests": "",
      "location": ""
    },
    "fulfillmentText": "I can help with that. Where would you like to reserve a room?",
    "fulfillmentMessages": [
      {
        "text": {
          "text": [
            "I can help with that. Where would you like to reserve a room?"
          ]
        }
      }
    ],
    "intent": {
      "name": "projects/PROJECT_ID/agent/intents/e8f6a63e-73da-4a1a-8bfc-857183f71228",
      "displayName": "room.reservation"
    },
    "intentDetectionConfidence": 1,
    "diagnosticInfo": {},
    "languageCode": "en-us"
  },
  "outputAudio": "UklGRs6vAgBXQVZFZm10IBAAAAABAAEAwF0AAIC7AA..."
}

שימו לב שהערך בשדה queryResult.action הוא room.reservation, והשדה outputAudio מכיל מחרוזת ארוכה של אודיו בפורמט Base64.

3. השמעת פלט האודיו

מעתיקים את הטקסט מהשדה outputAudio ושומרים אותו בקובץ בשם output_audio.b64. צריך להמיר את הקובץ הזה לאודיו.

דוגמה ל-Linux:

base64 -d output_audio.b64 > output_audio.wav

דוגמאות לפלטפורמות אחרות מופיעות במאמר פענוח תוכן אודיו בקידוד Base64 במאמרי העזרה של ה-API להמרת טקסט לדיבור.

עכשיו אפשר להפעיל את קובץ האודיו output_audio.wav ולשמוע שהוא תואם לטקסט בשדה queryResult.fulfillmentMessages[1].text.text[0] שלמעלה. האלמנט השני fulfillmentMessages נבחר כי הוא תגובת הטקסט לפלטפורמת ברירת המחדל.

Java

כדי לבצע אימות ב-Dialogflow CX, צריך להגדיר את Application Default Credentials. מידע נוסף זמין במאמר הגדרת אימות לסביבת פיתוח מקומית.


import com.google.api.gax.rpc.ApiException;
import com.google.cloud.dialogflow.v2.DetectIntentRequest;
import com.google.cloud.dialogflow.v2.DetectIntentResponse;
import com.google.cloud.dialogflow.v2.OutputAudioConfig;
import com.google.cloud.dialogflow.v2.OutputAudioEncoding;
import com.google.cloud.dialogflow.v2.QueryInput;
import com.google.cloud.dialogflow.v2.QueryResult;
import com.google.cloud.dialogflow.v2.SessionName;
import com.google.cloud.dialogflow.v2.SessionsClient;
import com.google.cloud.dialogflow.v2.TextInput;
import com.google.common.collect.Maps;
import java.io.IOException;
import java.util.List;
import java.util.Map;

public class DetectIntentWithTextToSpeechResponse {

  public static Map<String, QueryResult> detectIntentWithTexttoSpeech(
      String projectId, List<String> texts, String sessionId, String languageCode)
      throws IOException, ApiException {
    Map<String, QueryResult> queryResults = Maps.newHashMap();
    // Instantiates a client
    try (SessionsClient sessionsClient = SessionsClient.create()) {
      // Set the session name using the sessionId (UUID) and projectID (my-project-id)
      SessionName session = SessionName.of(projectId, sessionId);
      System.out.println("Session Path: " + session.toString());

      // Detect intents for each text input
      for (String text : texts) {
        // Set the text (hello) and language code (en-US) for the query
        TextInput.Builder textInput =
            TextInput.newBuilder().setText(text).setLanguageCode(languageCode);

        // Build the query with the TextInput
        QueryInput queryInput = QueryInput.newBuilder().setText(textInput).build();

        //
        OutputAudioEncoding audioEncoding = OutputAudioEncoding.OUTPUT_AUDIO_ENCODING_LINEAR_16;
        int sampleRateHertz = 16000;
        OutputAudioConfig outputAudioConfig =
            OutputAudioConfig.newBuilder()
                .setAudioEncoding(audioEncoding)
                .setSampleRateHertz(sampleRateHertz)
                .build();

        DetectIntentRequest dr =
            DetectIntentRequest.newBuilder()
                .setQueryInput(queryInput)
                .setOutputAudioConfig(outputAudioConfig)
                .setSession(session.toString())
                .build();

        // Performs the detect intent request
        DetectIntentResponse response = sessionsClient.detectIntent(dr);

        // Display the query result
        QueryResult queryResult = response.getQueryResult();

        System.out.println("====================");
        System.out.format("Query Text: '%s'\n", queryResult.getQueryText());
        System.out.format(
            "Detected Intent: %s (confidence: %f)\n",
            queryResult.getIntent().getDisplayName(), queryResult.getIntentDetectionConfidence());
        System.out.format(
            "Fulfillment Text: '%s'\n",
            queryResult.getFulfillmentMessagesCount() > 0
                ? queryResult.getFulfillmentMessages(0).getText()
                : "Triggered Default Fallback Intent");

        queryResults.put(text, queryResult);
      }
    }
    return queryResults;
  }
}

Node.js

כדי לבצע אימות ב-Dialogflow CX, צריך להגדיר את Application Default Credentials. מידע נוסף זמין במאמר הגדרת אימות לסביבת פיתוח מקומית.

// Imports the Dialogflow client library
const dialogflow = require('@google-cloud/dialogflow').v2;

// Instantiate a DialogFlow client.
const sessionClient = new dialogflow.SessionsClient();

/**
 * TODO(developer): Uncomment the following lines before running the sample.
 */
// const projectId = 'ID of GCP project associated with your Dialogflow agent';
// const sessionId = `user specific ID of session, e.g. 12345`;
// const query = `phrase(s) to pass to detect, e.g. I'd like to reserve a room for six people`;
// const languageCode = 'BCP-47 language code, e.g. en-US';
// const outputFile = `path for audio output file, e.g. ./resources/myOutput.wav`;

// Define session path
const sessionPath = sessionClient.projectAgentSessionPath(
  projectId,
  sessionId
);
const fs = require('fs');
const util = require('util');

async function detectIntentwithTTSResponse() {
  // The audio query request
  const request = {
    session: sessionPath,
    queryInput: {
      text: {
        text: query,
        languageCode: languageCode,
      },
    },
    outputAudioConfig: {
      audioEncoding: 'OUTPUT_AUDIO_ENCODING_LINEAR_16',
    },
  };
  sessionClient.detectIntent(request).then(responses => {
    console.log('Detected intent:');
    const audioFile = responses[0].outputAudio;
    util.promisify(fs.writeFile)(outputFile, audioFile, 'binary');
    console.log(`Audio content written to file: ${outputFile}`);
  });
}
detectIntentwithTTSResponse();

Python

כדי לבצע אימות ב-Dialogflow CX, צריך להגדיר את Application Default Credentials. מידע נוסף זמין במאמר הגדרת אימות לסביבת פיתוח מקומית.

def detect_intent_with_texttospeech_response(
    project_id, session_id, texts, language_code
):
    """Returns the result of detect intent with texts as inputs and includes
    the response in an audio format.

    Using the same `session_id` between requests allows continuation
    of the conversation."""
    from google.cloud import dialogflow

    session_client = dialogflow.SessionsClient()

    session_path = session_client.session_path(project_id, session_id)
    print("Session path: {}\n".format(session_path))

    for text in texts:
        text_input = dialogflow.TextInput(text=text, language_code=language_code)

        query_input = dialogflow.QueryInput(text=text_input)

        # Set the query parameters with sentiment analysis
        output_audio_config = dialogflow.OutputAudioConfig(
            audio_encoding=dialogflow.OutputAudioEncoding.OUTPUT_AUDIO_ENCODING_LINEAR_16
        )

        request = dialogflow.DetectIntentRequest(
            session=session_path,
            query_input=query_input,
            output_audio_config=output_audio_config,
        )
        response = session_client.detect_intent(request=request)

        print("=" * 20)
        print("Query text: {}".format(response.query_result.query_text))
        print(
            "Detected intent: {} (confidence: {})\n".format(
                response.query_result.intent.display_name,
                response.query_result.intent_detection_confidence,
            )
        )
        print("Fulfillment text: {}\n".format(response.query_result.fulfillment_text))
        # The response's audio_content is binary.
        with open("output.wav", "wb") as out:
            out.write(response.output_audio)
            print('Audio content written to file "output.wav"')

מידע נוסף על שדות התגובה הרלוונטיים מופיע בקטע תשובות לזיהוי כוונות.

זיהוי תגובות לכוונות

התגובה לבקשה לזיהוי כוונות היא מסוג DetectIntentResponse.

העיבוד הרגיל של זיהוי הכוונה שולט בתוכן של השדה DetectIntentResponse.queryResult.fulfillmentMessages.

השדה DetectIntentResponse.outputAudio מאוכלס באודיו על סמך הערכים של תגובות טקסט של פלטפורמת ברירת המחדל שנמצאות בשדה DetectIntentResponse.queryResult.fulfillmentMessages:

  • אם יש כמה תשובות טקסט שמוגדרות כברירת מחדל, הן יצורפו יחד כשיוצרים אודיו.
  • אם לא קיימות תשובות טקסט גנריות, התוכן האודיו שנוצר יהיה ריק.

השדה DetectIntentResponse.outputAudioConfig מאוכלס בהגדרות האודיו ששימשו ליצירת פלט האודיו.

זיהוי כוונות משידור

כשמזהים כוונה מסטרימינג, שולחים בקשות שדומות לדוגמה Detecting Intent from a Stream שלא משתמשת בפלט אודיו. עם זאת, אתם מספקים שדה OutputAudioConfig בבקשה. השדות output_audio ו-output_audio_config מאוכלסים בתגובת הסטרימינג הסופית שמתקבלת משרת Dialogflow API. מידע נוסף זמין במאמרים בנושא StreamingDetectIntentRequest ו-StreamingDetectIntentResponse.

הגדרות הסוכן לדיבור

אתם יכולים לשלוט בהיבטים שונים של סינתזת דיבור. הגדרות הדיבור של הסוכן

שימוש בסימולטור של Dialogflow

אתם יכולים לנהל אינטראקציה עם הסוכן ולקבל תשובות קוליות באמצעות הסימולטור של Dialogflow:

  1. כדי להפעיל המרת טקסט לדיבור באופן אוטומטי, פועלים לפי ההוראות במאמר בנושא הגדרות הדיבור של נציגים.
  2. מקלידים או אומרים 'book a room' (הזמנת חדר) בסימולטור.
  3. מעיינים בקטע פלט אודיו בסימולטור.