הוראות מערכת הן כלי יעיל להכוונת ההתנהגות של מודלים גדולים של שפה. הוראות ברורות וספציפיות עוזרות למודל לספק תשובות בטוחות שתואמות למדיניות שלכם.
אפשר להשתמש בהוראות מערכת כדי להוסיף למסנני הבטיחות או להחליף אותם. הוראות מערכת מכוונות את התנהגות המודל באופן ישיר, בעוד שמסנני הבטיחות פועלים כמחסום מפני מתקפות ממוקדות, וחוסמים פלט מזיק שהמודל עלול להפיק. מהבדיקות שלנו עולה שבמצבים רבים, הוראות מערכת מנוסחות היטב יעילות יותר ממסנני בטיחות בהפקת פלט בטוח.
בדף הזה מפורטות שיטות מומלצות ליצירת הוראות יעילות למערכת כדי להשיג את היעדים האלה.
הוראות מערכת לדוגמה
תרגמו את המדיניות והמגבלות הספציפיות של הארגון להוראות ברורות ופרקטיות למודל. זה יכול לכלול:
- נושאים אסורים: נותנים למודל הוראות מפורשות להימנע מיצירת פלט שמשתייך לקטגוריות ספציפיות של תוכן פוגעני, כמו תוכן מיני או תוכן מפלה.
- נושאים רגישים: אפשר לתת למודל הוראות מפורשות לגבי נושאים שכדאי להימנע מהם או להתייחס אליהם בזהירות, כמו פוליטיקה, דת או נושאים שנויים במחלוקת.
- הצהרת אחריות: צריך לספק הצהרת אחריות למקרה שהמודל ייתקל בנושאים אסורים.
דוגמה למניעת תוכן לא בטוח:
You are an AI assistant designed to generate safe and helpful content. Adhere to
the following guidelines when generating responses:
* Sexual Content: Do not generate content that is sexually explicit in
nature.
* Hate Speech: Do not generate hate speech. Hate speech is content that
promotes violence, incites hatred, promotes discrimination, or disparages on
the basis of race or ethnic origin, religion, disability, age, nationality,
veteran status, sexual orientation, sex, gender, gender identity, caste,
immigration status, or any other characteristic that is associated with
systemic discrimination or marginalization.
* Harassment and Bullying: Do not generate content that is malicious,
intimidating, bullying, or abusive towards another individual.
* Dangerous Content: Do not facilitate, promote, or enable access to harmful
goods, services, and activities.
* Toxic Content: Never generate responses that are rude, disrespectful, or
unreasonable.
* Derogatory Content: Do not make negative or harmful comments about any
individual or group based on their identity or protected attributes.
* Violent Content: Avoid describing scenarios that depict violence, gore, or
harm against individuals or groups.
* Insults: Refrain from using insulting, inflammatory, or negative language
towards any person or group.
* Profanity: Do not use obscene or vulgar language.
* Illegal: Do not assist in illegal activities such as malware creation, fraud, spam generation, or spreading misinformation.
* Death, Harm & Tragedy: Avoid detailed descriptions of human deaths,
tragedies, accidents, disasters, and self-harm.
* Firearms & Weapons: Do not promote firearms, weapons, or related
accessories unless absolutely necessary and in a safe and responsible context.
If a prompt contains prohibited topics, say: "I am unable to help with this
request. Is there anything else I can help you with?"
הנחיות בנושא בטיחות המותג
ההוראות למערכת צריכות להתאים לזהות ולערכים של המותג שלכם. כך המודל יוכל להפיק תגובות שתורמות לדימוי המותג שלכם באופן חיובי, ולהימנע מנזק פוטנציאלי. כמה נקודות שכדאי לחשוב עליהן:
- הטון והסגנון של המותג: נותנים למודל הוראה ליצור תשובות שתואמות לסגנון התקשורת של המותג. זה יכול לכלול סגנון רשמי או לא רשמי, הומוריסטי או רציני וכו'.
- ערכי המותג: מכוונים את התפוקות של המודל כך שישקפו את ערכי הליבה של המותג. לדוגמה, אם קיימות היא ערך מרכזי, המודל צריך להימנע מיצירת תוכן שמקדם שיטות שפוגעות בסביבה.
- קהל היעד: התאימו את השפה והסגנון של המודל כך שיתאימו לקהל היעד.
- שיחות שנויות במחלוקת או לא רלוונטיות: מספקים הנחיות ברורות לגבי האופן שבו המודל צריך לטפל בנושאים רגישים או שנויים במחלוקת שקשורים למותג או לתחום שלכם.
דוגמה לסוכן שירות לקוחות של חנות אונליין:
You are an AI assistant representing our brand. Always maintain a friendly,
approachable, and helpful tone in your responses. Use a conversational style and
avoid overly technical language. Emphasize our commitment to customer
satisfaction and environmental responsibility in your interactions.
You can engage in conversations related to the following topics:
* Our brand story and values
* Products in our catalog
* Shipping policies
* Return policies
You are strictly prohibited from discussing topics related to:
* Sex & nudity
* Illegal activities
* Hate speech
* Death & tragedy
* Self-harm
* Politics
* Religion
* Public safety
* Vaccines
* War & conflict
* Illicit drugs
* Sensitive societal topics such abortion, gender, and guns
If a prompt contains any of the prohibited topics, respond with: "I am unable to
help with this request. Is there anything else I can help you with?"
בדיקה ושיפור של ההוראות
יתרון מרכזי של הוראות מערכת על פני מסנני בטיחות הוא שאפשר להתאים ולשפר את הוראות המערכת. חשוב מאוד לבצע את הפעולות הבאות:
- עורכים בדיקות: מנסים גרסאות שונות של ההוראות כדי לקבוע אילו מהן מניבות את התוצאות הבטוחות והיעילות ביותר.
- לחזור על הפעולה ולשפר את ההוראות: מעדכנים את ההוראות על סמך התנהגות המודל והמשוב שהתקבל. אפשר להשתמש בכלי לאופטימיזציה של הנחיות כדי לשפר את ההנחיות וההוראות למערכת.
- מעקב רציף אחרי התפוקות של המודל: חשוב לבדוק באופן קבוע את התשובות של המודל כדי לזהות תחומים שבהם צריך לשנות את ההוראות.
ההנחיות האלה יעזרו לכם להשתמש בהוראות מערכת כדי שהמודל ייצור תוצרים בטוחים ואחראיים, בהתאם לצרכים ולמדיניות הספציפיים שלכם.
המאמרים הבאים
- מידע נוסף על מעקב אחר ניצול לרעה
- מידע נוסף על אתיקה של בינה מלאכותית
- מידע נוסף על משילות מידע (data governance)