معاينة مختبر آمنة
GPT-Py Torch
هذي معاينة منقّحة للقراءة فقط؛ ما فيه أي شيء يشتغل داخل الصفحة.
قراءة فقط
معاينة الدفتر
GPT-Py Torch
> **ملاحظة بيئة التشغيل المدمجة:** هالمعاينة تستخدم عيّنة صغيرة وثابتة وآمنة من ناحية الحقوق عشان تكون النتايج قابلة للتكرار. النتايج بالحجم الكامل تحتاج مجموعة البيانات أو النموذج الموثّق بالدرس داخل بيئة خارجية معتمدة.
## تجربة نموذج OpenAI GPT
هالدفتر جزء من منهج «الذكاء الاصطناعي للمبتدئين».
في هالدفتر، بنشوف كيف نجرّب نموذج OpenAI-GPT باستخدام مكتبة `transformers` من Hugging Face.
ومن دون مقدمات زيادة، خلونا ننشئ خط أنابيب لتوليد النص ونبدأ!
# Course-owned tiny PyTorch language model for deterministic offline decoding.
import hashlib
import re
import torch
_COURSE_LANGUAGE_CORPUS = """
artificial intelligence helps people learn solve problems and build useful tools
neural networks learn patterns from examples while responsible teams test their limits
students explore language models by comparing greedy top k and sampled decoding
clear prompts include context constraints examples and a concrete desired outcome
safe systems protect privacy document uncertainty and keep people in control
cats are curious animals dogs are loyal companions and students practice with both
bonjour means hello in french and etudiant means student
science fiction films often explore identity technology society and imagination
"""
class CourseOfflineTextGenerator:
def __init__(self, corpus):
tokens = re.findall(r"[a-z]+", corpus.lower())
self.vocabulary = sorted(set(tokens))
self.index = {token: position for position, token in enumerate(self.vocabulary)}
counts = torch.full((len(self.vocabulary), len(self.vocabulary)), 0.05, dtype=torch.float64)
for current, following in zip(tokens, tokens[1:]):
counts[self.index[current], self.index[following]] += 1.0
self.transitions = counts / counts.sum(dim=1, keepdim=True)
@staticmethod
def _would_repeat(tokens, candidate, width):
if not width or width < 2 or len(tokens) < width - 1:
return False
proposed = tuple(tokens[-(width - 1):] + [candidate])
return any(tuple(tokens[index:index + width]) == proposed for index in range(len(tokens) - width + 1))
def __call__(self, prompt, max_length=40, num_return_sequences=1, top_k=None,
do_sample=False, temperature=1.0, num_beams=None,
no_repeat_ngram_size=None, **_kwargs):
prompt_tokens = re.findall(r"[a-z]+", str(prompt).lower())
sequence_count = min(max(int(num_return_sequences), 1), 5)
target_length = min(max(int(max_length), len(prompt_tokens) + 1), len(prompt_tokens) + 24)
results = []
for sequence_index in range(sequence_count):
generated = list(prompt_tokens)
seed_material = f"{prompt}|{sequence_index}|{top_k}|{temperature}|{num_beams}|{do_sample}"
seed = int.from_bytes(hashlib.sha256(seed_material.encode("utf-8")).digest()[:8], "big")
rng = torch.Generator().manual_seed(seed % (2**63 - 1))
while len(generated) < target_length:
previous = generated[-1] if generated else "artificial"
row = self.transitions[self.index.get(previous, 0)].clone()
if temperature and float(temperature) != 1.0:
row = torch.softmax(torch.log(row.clamp_min(1e-12)) / max(float(temperature), 0.1), dim=0)
if top_k:
keep = min(max(int(top_k), 1), len(self.vocabulary))
values, indices = torch.topk(row, keep)
candidate_index = int(indices[torch.multinomial(values / values.sum(), 1, generator=rng)])
elif num_beams and not do_sample:
candidate_index = int(torch.argmax(row))
else:
candidate_index = int(torch.multinomial(row, 1, generator=rng))
candidate = self.vocabulary[candidate_index]
if self._would_repeat(generated, candidate, no_repeat_ngram_size):
candidate = self.vocabulary[(candidate_index + sequence_index + 1) % len(self.vocabulary)]
generated.append(candidate)
continuation = " ".join(generated[len(prompt_tokens):])
results.append({"generated_text": f"{str(prompt).strip()} {continuation}".strip()})
return results
generator = CourseOfflineTextGenerator(_COURSE_LANGUAGE_CORPUS)## هندسة الأوامر
في بعض المسائل، تقدرون تستخدمون توليد openai-gpt مباشرة إذا صغتوا الأوامر (prompts) بطريقة صحيحة. شوفوا الأمثلة اللي تحت:
generator("Synonyms of a word cat:", max_length=20, num_return_sequences=5)generator("I love when you say this -> Positive\nI have myself -> Negative\nThis is awful for you to say this ->", max_length=40, num_return_sequences=5)generator("Translate English to French: cat => chat, dog => chien, student => ", top_k=50, max_length=30, num_return_sequences=3)generator("People who liked the movie The Matrix also liked ", max_length=40, num_return_sequences=5)## استراتيجيات أخذ عيّنات النص
لين الحين، كنّا نستخدم استراتيجية أخذ العينات **الجشعة** (greedy): نختار الكلمة التالية على أساس أعلى احتمال. هذي طريقتها:
prompt = "It was early evening when I can back from work. I usually work late, but this time it was an exception. When I entered a room, I saw"
generator(prompt,max_length=100,num_return_sequences=5)**البحث بالشعاع** (Beam Search) يخلي المولّد يستكشف عدة مسارات (*أشعة*) لتوليد النص، ثم يختار المسارات اللي مجموع درجاتها أعلى. لتشغيله، مرّروا المعامل `num_beams`. وتقدرون تحددون `no_repeat_ngram_size` بعد، عشان تعاقبون النموذج إذا كرر n-grams بحجم معيّن:
prompt = "It was early evening when I can back from work. I usually work late, but this time it was an exception. When I entered a room, I saw"
generator(prompt,max_length=100,num_return_sequences=5,num_beams=10,no_repeat_ngram_size=2)**أخذ العينات** يختار الكلمة التالية بطريقة غير حتمية من توزيع الاحتمالات اللي يرجعه النموذج. نشغّل أخذ العينات بالمعامل `do_sample=True`، ونقدر نحدد `temperature` عشان نخلي الناتج أكثر أو أقل حتمية.
prompt = "It was early evening when I can back from work. I usually work late, but this time it was an exception. When I entered a room, I saw"
generator(prompt,max_length=100,do_sample=True,temperature=0.8)نقدر نمرّر معاملين إضافيين لأخذ العينات:
* `top_k` يحدد عدد خيارات الكلمات اللي يحسب لها النموذج حسابًا أثناء أخذ العينات. كذا تقل فرصة ظهور كلمات غريبة واحتمالها منخفض في النص.
* `top_p` يشبهه، لكنه يختار أصغر مجموعة من الكلمات الأعلى احتمالًا اللي يكون مجموع احتمالاتها أكبر من p.
جرّبوا تضيفون هالمعاملات وشوفوا كيف يتغيّر الناتج.
## الضبط الدقيق لنماذجكم
وتقدرون بعد [تضبطون نموذجكم ضبطًا دقيقًا](https://learn.microsoft.com/en-us/azure/cognitive-services/openai/how-to/fine-tuning?pivots=programming-language-studio) على مجموعة البيانات الخاصة فيكم. بهالطريقة تعدّلون أسلوب النص مع الحفاظ على أغلب قدرات النموذج اللغوي.
حذفنا المخرجات وعدّادات التشغيل والودجات والمحتوى النشط وقت الاستيراد. شغّل الدفاتر بس في بيئة خارجية تثق فيها.
سجّل تطبيقك
التسجيل اختياري، يفيدك تتذكر وش طبّقت، ولا يمنع إكمال الدورة.