معاينة مختبر آمنة
Text Representation Py Torch
هذي معاينة منقّحة للقراءة فقط؛ ما فيه أي شيء يشتغل داخل الصفحة.
قراءة فقط
معاينة الدفتر
Text Representation Py Torch
> **ملاحظة بيئة التشغيل المدمجة:** هالمعاينة تستخدم عيّنة صغيرة وثابتة وآمنة من ناحية الحقوق عشان تكون النتايج قابلة للتكرار. النتايج بالحجم الكامل تحتاج مجموعة البيانات أو النموذج الموثّق بالدرس داخل بيئة خارجية معتمدة.
# مهمة تصنيف النصوص
بنركز على مهمة بسيطة في تصنيف النصوص باستخدام مجموعة أخبار صغيرة مكتوبة خصيصًا لهالنسخة. المطلوب نصنّف كل خبر ضمن أربع فئات: العالم، والرياضة، والأعمال، والعلوم/التقنية.
## مجموعة البيانات
مجموعة البيانات مضمّنة داخل الدفتر، عشان تكون التجربة ثابتة وسريعة وما تعتمد على تنزيلات أو مكتبات بيانات متوقفة.
import collections
import random
import re
import numpy as np
import torch
# course-edition deterministic NLP utilities v2
COURSE_NEWS_CLASSES = ["World", "Sports", "Business", "Sci/Tech"]
# This compact corpus was written for this edition. It keeps the notebooks
# deterministic, network-free, and quick enough to run on a learner's CPU.
COURSE_NEWS_TRAIN = [
(1, "Paris hosted a regional summit where diplomats discussed water security, peaceful trade routes, and a shared emergency plan for neighboring communities."),
(1, "The elected council approved a cross-border health agreement after delegates reviewed hospital capacity, medicine access, and transparent public reporting."),
(1, "A coastal city welcomed observers for a public vote, while local groups published clear guidance about polling places and accessible transportation."),
(1, "The king and queen met a woman who leads the relief agency and a man who coordinates volunteers, then the group announced a neutral humanitarian corridor."),
(1, "Ministers signed a climate adaptation pledge that funds drought monitoring, resilient farms, and open scientific exchange across several countries."),
(1, "Community mediators opened a week of talks focused on civilian safety, reliable food deliveries, and practical steps toward a lasting ceasefire."),
(2, "The basketball team completed a patient comeback in the final quarter, using quick passes, strong defense, and a balanced scoring plan to win the tournament."),
(2, "A young runner broke the course record after months of careful training, recovery sessions, nutrition planning, and support from her local athletics club."),
(2, "The football coach praised disciplined play after the squad protected an early lead and created chances through accurate movement on both wings."),
(2, "Fans filled the arena for a close volleyball final in which both teams served aggressively, defended long rallies, and respected every referee decision."),
(2, "The tennis champion returned after injury and won a demanding match by varying pace, placing serves carefully, and staying calm during two tie breaks."),
(2, "Organizers added inclusive swimming events to the city games so more athletes can compete with safe facilities, trained officials, and fair timing equipment."),
(3, "Microsoft announced a small-business program that combines cloud credits, security workshops, and practical accounting support for new local companies."),
(3, "Global funds moved cautiously after the central bank held interest rates steady and asked lenders to publish clearer information about consumer borrowing costs."),
(3, "A neighborhood market expanded its delivery service, hired twelve workers, and invested revenue in reusable packaging supplied by another local business."),
(3, "The manufacturer reported stable quarterly sales while noting that shipping delays and higher material prices could reduce margins later in the year."),
(3, "Two payment companies agreed to share fraud signals through a privacy-preserving system designed to protect customers without exposing purchase details."),
(3, "A cooperative offered farmers transparent contracts and faster invoices, helping members plan equipment repairs before the next growing season begins."),
(4, "Researchers trained a compact neural network to identify damaged solar panels, then published the evaluation data and documented where the model still fails."),
(4, "A university technology lab released an open sensor design that measures classroom air quality and stores only anonymous readings for public analysis."),
(4, "Engineers improved a battery recycling process by recovering more useful minerals at lower temperatures, reducing energy use during the pilot study."),
(4, "The space telescope captured a detailed spectrum from a distant planet, giving scientists new evidence about clouds and molecules in its atmosphere."),
(4, "A software team tested an accessibility assistant with keyboard users and screen-reader experts before changing the interface and publishing the results."),
(4, "Students built a small robot that maps indoor obstacles with inexpensive sensors, explains each route choice, and says hello when a test run begins."),
]
COURSE_NEWS_TEST = [
(1, "Regional delegates published a joint disaster response schedule after reviewing evacuation routes and communication gaps."),
(1, "Election monitors confirmed the final count and recommended clearer access rules for voters who need assistance."),
(2, "The basketball captain scored late, but credited the victory to defense, passing, and careful preparation across the whole season."),
(2, "A cycling club opened a safe youth race with trained marshals, marked turns, and free equipment checks."),
(3, "Technology shares rose while retail funds remained cautious after companies issued mixed forecasts for the next quarter."),
(3, "The family business secured a modest loan to replace old equipment and expand its apprenticeship program."),
(4, "A research team released a smaller language model with documented energy measurements and a public evaluation set."),
(4, "Scientists used open satellite data to improve flood warnings and shared the software with local emergency teams."),
]
def course_tokenize(text):
return re.findall(r"[a-z0-9]+(?:'[a-z0-9]+)?", str(text).lower())
def course_ngrams_iterator(tokens, ngrams=1):
tokens = list(tokens)
for token in tokens:
yield token
for size in range(2, max(1, ngrams) + 1):
for start in range(0, len(tokens) - size + 1):
yield " ".join(tokens[start:start + size])
class CourseVocabulary:
def __init__(self, counter, min_freq=1, max_tokens=None):
ordered = sorted(
(token for token, count in counter.items() if count >= min_freq),
key=lambda token: (-counter[token], token),
)
if max_tokens is not None:
ordered = ordered[:max(0, max_tokens - 1)]
self.itos = ["<unk>"] + [token for token in ordered if token != "<unk>"]
self.stoi = {token: index for index, token in enumerate(self.itos)}
def __len__(self):
return len(self.itos)
def __getitem__(self, token):
return self.stoi.get(token, 0)
def get_stoi(self):
return dict(self.stoi)
def get_itos(self):
return list(self.itos)
def build_course_vocab(dataset, ngrams=1, min_freq=1, max_tokens=None, tokenizer=course_tokenize):
counter = collections.Counter()
for _label, text in dataset:
counter.update(course_ngrams_iterator(tokenizer(text), ngrams=ngrams))
return CourseVocabulary(counter, min_freq=min_freq, max_tokens=max_tokens)
def load_course_fixture(ngrams=1, min_freq=1, vocab_size=None, lines_cnt=None):
global vocab, tokenizer
tokenizer = course_tokenize
train_dataset = list(COURSE_NEWS_TRAIN)
test_dataset = list(COURSE_NEWS_TEST)
vocabulary_rows = train_dataset if lines_cnt is None else train_dataset[:lines_cnt]
vocab = build_course_vocab(
vocabulary_rows,
ngrams=ngrams,
min_freq=min_freq,
max_tokens=vocab_size,
tokenizer=tokenizer,
)
return train_dataset, test_dataset, list(COURSE_NEWS_CLASSES), vocab
def encode(text, voc=None, unk=0, tokenizer=course_tokenize):
selected_vocab = vocab if voc is None else voc
stoi = selected_vocab.get_stoi()
return [stoi.get(token, unk) for token in tokenizer(text)]
def train_epoch(net, dataloader, lr=0.01, optimizer=None, loss_fn=None, epoch_size=None, report_freq=200):
optimizer = optimizer or torch.optim.Adam(net.parameters(), lr=lr)
loss_fn = (loss_fn or torch.nn.CrossEntropyLoss()).to(device)
net.train()
total_loss, accuracy, count, batch_index = 0.0, 0, 0, 0
for labels, features in dataloader:
optimizer.zero_grad()
features, labels = features.to(device), labels.to(device)
output = net(features)
loss = loss_fn(output, labels)
loss.backward()
optimizer.step()
total_loss += loss.item()
accuracy += (output.argmax(1) == labels).sum().item()
count += len(labels)
batch_index += 1
if batch_index % report_freq == 0:
print(f"{count}: acc={accuracy / count:.4f}")
if epoch_size and count > epoch_size:
break
return total_loss / max(count, 1), accuracy / max(count, 1)
def padify(batch, voc=None, tokenizer=course_tokenize):
vectors = [encode(item[1], voc=voc, tokenizer=tokenizer) for item in batch]
max_length = max(map(len, vectors))
return (
torch.LongTensor([item[0] - 1 for item in batch]),
torch.stack([
torch.nn.functional.pad(torch.tensor(vector), (0, max_length - len(vector)), value=0)
for vector in vectors
]),
)
def offsetify(batch, voc=None):
vectors = [torch.tensor(encode(item[1], voc=voc)) for item in batch]
offsets = torch.tensor(([0] + [len(vector) for vector in vectors])[:-1]).cumsum(dim=0)
return torch.LongTensor([item[0] - 1 for item in batch]), torch.cat(vectors), offsets
def train_epoch_emb(net, dataloader, lr=0.01, optimizer=None, loss_fn=None, epoch_size=None, report_freq=200, use_pack_sequence=False):
optimizer = optimizer or torch.optim.Adam(net.parameters(), lr=lr)
loss_fn = (loss_fn or torch.nn.CrossEntropyLoss()).to(device)
net.train()
total_loss, accuracy, count, batch_index = 0.0, 0, 0, 0
for labels, text, offsets in dataloader:
optimizer.zero_grad()
labels, text = labels.to(device), text.to(device)
offsets = offsets.to("cpu" if use_pack_sequence else device)
output = net(text, offsets)
loss = loss_fn(output, labels)
loss.backward()
optimizer.step()
total_loss += loss.item()
accuracy += (output.argmax(1) == labels).sum().item()
count += len(labels)
batch_index += 1
if batch_index % report_freq == 0:
print(f"{count}: acc={accuracy / count:.4f}")
if epoch_size and count > epoch_size:
break
return total_loss / max(count, 1), accuracy / max(count, 1)
random.seed(42)
np.random.seed(42)
torch.manual_seed(42)
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
tokenizer = course_tokenize
vocab = None
train_dataset, test_dataset, classes, vocab = load_course_fixture()هنا، يحتوي `train_dataset` و`test_dataset` على مجموعتين ترجع كل وحدة منها زوجًا: التسمية (رقم الفئة) ثم النص، مثل كذا:
list(train_dataset)[0]خلونا الحين نطبع أول 10 عناوين أخبار من مجموعة البيانات:
for i, x in zip(range(5), train_dataset):
print(f"**{classes[x[0] - 1]}** -> {x[1]}")مجموعات البيانات هنا مكرّرات (iterators)، فإذا بغينا نستخدم البيانات أكثر من مرة لازم نحوّلها إلى قائمة:
train_dataset, test_dataset, classes, vocab = load_course_fixture()## تجزئة النص إلى رموز
الحين نحتاج نحوّل النص إلى **أرقام** نقدر نمثّلها على هيئة موترات. وإذا بغينا تمثيلًا على مستوى الكلمات، عندنا خطوتان:
* نستخدم **مجزّئ نص** يقسّم النص إلى **رموز**.
* نبني **مفردات** من هالرموز.
tokenizer('He said: hello')counter = collections.Counter()
for label, line in train_dataset:
counter.update(tokenizer(line))
vocab = CourseVocabulary(counter, min_freq=1)وبالاعتماد على المفردات، نقدر نحوّل سلسلة الرموز بسهولة إلى مجموعة أرقام:
vocab_size = len(vocab)
print(f"Vocab size is {vocab_size}")
stoi = vocab.get_stoi()
def encode(x):
return [stoi.get(token, 0) for token in tokenizer(x)]
encode('I love to play with my words')## تمثيل النص بنموذج حقيبة الكلمات
بما إن الكلمات تحمل معنى، نقدر أحيانًا نستنتج موضوع النص من الكلمات الموجودة فيه لحالها، حتى لو تجاهلنا ترتيبها في الجملة. مثلًا، في تصنيف الأخبار، كلمتا *weather* و*snow* غالبًا تدلّان على *weather forecast*، بينما *stocks* و*dollar* تميلان إلى الأخبار المالية.
**نموذج حقيبة الكلمات** (BoW) هو أكثر تمثيلات المتجهات التقليدية استخدامًا. نربط كل كلمة بفهرس داخل المتجه، وتكون قيمة العنصر هي عدد مرات ظهور هالكلمة في مستند معيّن.
> **وصف الشكل:** رسم يوضّح كيف يُخزّن تمثيل متجه حقيبة الكلمات في الذاكرة.
> **ملاحظة:** نقدر بعد نفهم BoW على إنه مجموع متجهات one-hot لكل كلمة في النص.
وهذا مثال يوضّح كيف نولّد تمثيل حقيبة الكلمات باستخدام مكتبة Scikit Learn في Python:
from sklearn.feature_extraction.text import CountVectorizer
vectorizer = CountVectorizer()
corpus = [
'I like hot dogs.',
'The dog ran fast.',
'Its hot outside.',
]
vectorizer.fit_transform(corpus)
vectorizer.transform(['My dog likes hot dogs on a hot day.']).toarray()ولحساب متجه حقيبة الكلمات من التمثيل المتجهي لمجموعة البيانات AG_NEWS، نستخدم الدالة التالية:
vocab_size = len(vocab)
def to_bow(text,bow_vocab_size=vocab_size):
res = torch.zeros(bow_vocab_size,dtype=torch.float32)
for i in encode(text):
if i<bow_vocab_size:
res[i] += 1
return res
print(to_bow(train_dataset[0][1]))> **ملاحظة:** نستخدم هنا المتغير العام `vocab_size` لتحديد الحجم الافتراضي للمفردات. وبما إن المفردات غالبًا كبيرة، نقدر نحصرها في الكلمات الأكثر تكرارًا. جرّبوا تقلّلون قيمة `vocab_size` وتشغّلون الكود اللي تحت، ثم شوفوا وش يصير على الدقة. المتوقع تنزل الدقة شوي، مب نزولًا كبيرًا، مقابل أداء أفضل.
## تدريب مصنّف BoW
بعد ما عرفنا كيف نبني تمثيل نموذج حقيبة الكلمات للنص، خلونا ندرّب مصنّفًا فوقه. أولًا، نحتاج نجهّز مجموعة البيانات للتدريب بحيث تتحوّل كل التمثيلات المتجهية الموضعية إلى تمثيل حقيبة الكلمات. نسوي هالشي بتمرير الدالة `bowify` في المعامل `collate_fn` إلى `DataLoader` القياسي في torch:
from torch.utils.data import DataLoader
import numpy as np
# this collate function gets list of batch_size tuples, and needs to
# return a pair of label-feature tensors for the whole minibatch
def bowify(b):
return (
torch.LongTensor([t[0]-1 for t in b]),
torch.stack([to_bow(t[1]) for t in b])
)
train_loader = DataLoader(train_dataset, batch_size=16, collate_fn=bowify, shuffle=True)
test_loader = DataLoader(test_dataset, batch_size=16, collate_fn=bowify, shuffle=True)الحين بنعرّف **الشبكة العصبية** البسيطة المستخدمة في **التصنيف**، وفيها طبقة خطية وحدة. حجم متجه الإدخال يساوي `vocab_size`، وحجم الإخراج يساوي عدد الفئات (4). وبما إن المهمة تصنيف، تكون دالة التنشيط الأخيرة `LogSoftmax`.
net = torch.nn.Sequential(torch.nn.Linear(vocab_size,4),torch.nn.LogSoftmax(dim=1))الحين بنعرّف حلقة التدريب القياسية في PyTorch. مجموعة البيانات كبيرة، ولغرض الشرح بندرّب حقبة وحدة فقط، وأحيانًا أقل من حقبة؛ لأن المعامل `epoch_size` يخلينا نحدّ مدة التدريب. وبنطبع بعد الدقة التراكمية أثناء التدريب، ويحدّد المعامل `report_freq` كم مرة يطلع التقرير.
def train_epoch(net,dataloader,lr=0.01,optimizer=None,loss_fn = torch.nn.NLLLoss(),epoch_size=None, report_freq=200):
optimizer = optimizer or torch.optim.Adam(net.parameters(),lr=lr)
net.train()
total_loss,acc,count,i = 0,0,0,0
for labels,features in dataloader:
optimizer.zero_grad()
out = net(features)
loss = loss_fn(out,labels) #cross_entropy(out,labels)
loss.backward()
optimizer.step()
total_loss+=loss
_,predicted = torch.max(out,1)
acc+=(predicted==labels).sum()
count+=len(labels)
i+=1
if i%report_freq==0:
print(f"{count}: acc={acc.item()/count}")
if epoch_size and count>epoch_size:
break
return total_loss.item()/count, acc.item()/counttrain_epoch(net,train_loader,epoch_size=15000)## ثنائيات الكلمات وثلاثيات الكلمات وN-Grams
من قيود نموذج حقيبة الكلمات إن بعض الكلمات تجي ضمن تعبير مكوّن من أكثر من كلمة. مثلًا، عبارة 'hot dog' معناها يختلف تمامًا عن معنى 'hot' و'dog' كل وحدة في سياق ثاني. وإذا مثّلنا 'hot' و'dog' دائمًا بالمتجهات نفسها، ممكن نربك النموذج.
عشان نعالج هالمشكلة، تُستخدم **تمثيلات N-gram** كثير في **التصنيف** للمستندات؛ لأن تكرار الكلمة، أو زوج الكلمات، أو ثلاث كلمات متتابعة يكون سمة مفيدة لتدريب المصنّف. في تمثيل bigram مثلًا، نضيف كل أزواج الكلمات إلى المفردات بجانب الكلمات الأصلية.
وهذا مثال على توليد تمثيل حقيبة كلمات من نوع bigram باستخدام Scikit Learn:
bigram_vectorizer = CountVectorizer(ngram_range=(1, 2), token_pattern=r'\b\w+\b', min_df=1)
corpus = [
'I like hot dogs.',
'The dog ran fast.',
'Its hot outside.',
]
bigram_vectorizer.fit_transform(corpus)
print("Vocabulary:\n",bigram_vectorizer.vocabulary_)
bigram_vectorizer.transform(['My dog likes hot dogs on a hot day.']).toarray()العيب الأهم في أسلوب N-gram إن حجم المفردات يكبر بسرعة شديدة. عمليًا، نحتاج نجمع تمثيل N-gram مع تقنية لتقليل الأبعاد، مثل *التضمينات*، وبنتكلم عنها في الوحدة الجاية.
ولاستخدام تمثيل N-gram في مجموعة البيانات **AG News**، نحتاج نبني مفردات خاصة لـngram:
counter = collections.Counter()
for label, line in train_dataset:
counter.update(course_ngrams_iterator(tokenizer(line), ngrams=2))
bi_vocab = CourseVocabulary(counter, min_freq=1)
print("Bigram vocabulary length = ", len(bi_vocab))نقدر نستخدم الكود نفسه اللي فوق لتدريب المصنّف، لكنه بيستهلك الذاكرة بكفاءة سيئة. في الوحدة الجاية بندرّب مصنّف bigram باستخدام التضمينات.
> **ملاحظة:** تقدرون تخلّون بس ngrams اللي تظهر في النص أكثر من عدد محدد. كذا نستبعد أزواج الكلمات النادرة ونقلّل الأبعاد بشكل واضح. ارفعوا قيمة المعامل `min_freq`، وراقبوا كيف يتغيّر طول المفردات.
## تكرار المصطلح–معكوس تكرار المستند (TF-IDF)
في تمثيل BoW، تأخذ كل كلمة الوزن بالطريقة نفسها مهما كانت الكلمة. لكن واضح إن الكلمات الشائعة، مثل *a* و*in*، أقل فائدة في **التصنيف** من المصطلحات المتخصصة. وفي أغلب مهام معالجة اللغة الطبيعية (NLP) تكون بعض الكلمات أهم من غيرها.
يرمز **TF-IDF** إلى **تكرار المصطلح–معكوس تكرار المستند**. وهو تعديل على نموذج حقيبة الكلمات: بدل قيمة ثنائية 0/1 تدل على وجود الكلمة في المستند، نستخدم قيمة ذات فاصلة عائمة مرتبطة بمعدل ظهور الكلمة في المتن النصي.
وبصياغة أدق، نعرّف وزن الكلمة $i$ في المستند $j$، وهو $w_{ij}$، كالتالي:
$$
w_{ij} = tf_{ij}\times\log({N\over df_i})
$$
بحيث:
* $tf_{ij}$ هو عدد مرات ظهور $i$ في $j$؛ يعني قيمة BoW اللي شفناها قبل
* $N$ هو عدد المستندات في المجموعة
* $df_i$ هو عدد المستندات في المجموعة كلها اللي تحتوي الكلمة $i$
ترتفع قيمة TF-IDF، أي $w_{ij}$، طرديًا مع عدد مرات ظهور الكلمة في المستند، ويوازنها عدد المستندات اللي تحتوي هالكلمة في المتن النصي. بهالطريقة نعالج كون بعض الكلمات أكثر شيوعًا من غيرها. مثلًا، لو ظهرت الكلمة في *كل* مستندات المجموعة، بيكون $df_i=N$ و$w_{ij}=0$، ووقتها نتجاهل هالمصطلح بالكامل.
ونقدر بسهولة نسوي تمثيل TF-IDF للنص باستخدام Scikit Learn:
from sklearn.feature_extraction.text import TfidfVectorizer
vectorizer = TfidfVectorizer(ngram_range=(1,2))
vectorizer.fit_transform(corpus)
vectorizer.transform(['My dog likes hot dogs on a hot day.']).toarray()## الخلاصة
مع إن تمثيلات TF-IDF تعطي الكلمات أوزانًا حسب تكرارها، إلا إنها ما تمثّل المعنى ولا الترتيب. وكما قال اللغوي المعروف J. R. Firth سنة 1935: «المعنى الكامل للكلمة يعتمد دومًا على السياق، وما يصح نأخذ دراسة للمعنى بجدية إذا فصلناها عن سياقها». قدّام في الدورة بنتعلم كيف نلتقط معلومات السياق من النص باستخدام نمذجة اللغة.
حذفنا المخرجات وعدّادات التشغيل والودجات والمحتوى النشط وقت الاستيراد. شغّل الدفاتر بس في بيئة خارجية تثق فيها.
سجّل تطبيقك
التسجيل اختياري، يفيدك تتذكر وش طبّقت، ولا يمنع إكمال الدورة.