Evaluasi kinerja alur aplikasi Anda sebelum diterapkan di lingkungan produksi. LangStudio memberi skor aplikasi berdasarkan beberapa dimensi menggunakan templat evaluasi preset atau kustom.
Ikhtisar
Untuk mengevaluasi alur aplikasi, konfigurasikan set data evaluasi, petakan bidang input, dan pilih templat evaluasi. Sistem memproses setiap baris dataset secara batch melalui aplikasi Anda, memberi skor pada setiap output terhadap bidang tambahan, lalu mengagregasi hasilnya.

Sebelum memulai
-
Buat dan debug alur aplikasi. Kembangkan alur aplikasi.
-
Siapkan set data evaluasi dalam format JSONL dan unggah ke OSS. Contoh:
{"history":[],"query": "Describe the perilous majesty of Mount Hua", "reference": "Mount Hua stands alone, soaring to the clouds; \nSheer cliffs cut the sky, with rugged, handsome crags. \nGreen pines and bamboo vie for beauty on the cliffs; \nMonkeys cry and eagles fly, lit by frosty swords of light. \n\nPerilous peaks like scissors, jagged swords pointing to the sky; \nNarrow paths on steep slopes, where vines are the only way. \nWind and mist intertwine, as clouds emerge from caves; \nA deep fairyland, with a heavenly ladder hard to climb. \n\nJagged ridges cross, like a surging dragon's spine; \nDangerous paths lead onward, twisting toward the heavens. \nFrom lonely pine tops, eagles strike the vast sky; \nAt the summit of Mount Hua, a majestic and heroic sight.", "contexts": ["Mount Hua is one of the Five Great Mountains of China.", "Mount Hua is famous for its precipitous cliffs."]} {"history":[],"query": "Can you list 5 rare metals? Please rank them by global demand.", "reference": "Rare metals are metallic elements that are scarce in the Earth's crust, unevenly distributed, or difficult to mine. They play a crucial role in high-tech fields and emerging industries. The ranking of global demand can change with time and technological progress, but the following are some rare metals that are typically in high demand. This list is not necessarily ranked by absolute demand, as that can vary at different times.\n\n1. **Cobalt (Co)** - Cobalt is a key component of lithium-ion batteries, especially in electric vehicles and portable electronics. It is also used to manufacture heat-resistant alloys, hard alloys, and catalysts.\n\n2. **Neodymium (Nd)** - Neodymium is a rare-earth metal mainly used to produce strong magnets, such as high-performance permanent magnets. These magnets are widely used in computer hard drives, wind turbines, and the drive motors of electric vehicles.\n\n3. **Lithium (Li)** - Lithium is primarily used to manufacture lithium batteries. As the demand for electric vehicles and portable electronic devices increases, the demand for lithium is rising rapidly.\n\n4. **Silver (Ag)** - Although silver is not as rare as the metals listed above, its industrial demand is huge. It is mainly used in electronics, solar panels, jewelry, and currency manufacturing.\n\n5. **Ruthenium (Ru)** - Ruthenium is a rare precious metal widely used for data storage in hard disk drives and large-capacity servers. It is also used in catalysts and electrochemical cells.\n\nThe demand for these metals is influenced by many factors, such as the global economy, technological development, and policy support. Moreover, as time passes and markets change, other rare metals such as tantalum, indium, rhenium, and other rare-earth metals may also appear on the list of most in-demand rare metals.", "contexts": ["Rare metals are metals with low abundance in the Earth's crust that are complex to mine and extract.", "Lithium (Li): Used in battery manufacturing.", "Cobalt (Co): Used in high-performance alloys and battery manufacturing."]}File contoh: langstudio_eval_demo.jsonl
-
Buat koneksi LLM yang diperlukan untuk evaluasi. Konfigurasi koneksi.
Catatan: Beberapa templat evaluasi memerlukan model LLM-as-a-Judge. Konfigurasikan koneksi LLM yang sesuai sebelum menggunakan templat tersebut.
Penagihan
Evaluasi alur aplikasi menggunakan OSS untuk penyimpanan dataset dan PAI-DLC untuk menjalankan tugas offline. Anda dikenai biaya untuk kedua resource tersebut. Penagihan OSS. Penagihan Deep Learning Containers (DLC).
Buat tugas evaluasi
Setelah men-debug aplikasi Anda, klik Evaluation di pojok kanan atas untuk membuat tugas evaluasi.

Parameter utama:
|
Parameter |
Deskripsi |
|
Evaluation dataset |
|
|
OSS file |
Pilih file JSONL dari OSS sebagai set data evaluasi Anda. Dataset harus berisi bidang |
|
Application flow input mapping |
|
|
question/chat_history |
Petakan bidang input aplikasi ke kolom dataset. Catatan: Tugas evaluasi menjalankan aplikasi Anda untuk inferensi sebelum memberi skor hasil. Pilih bidang input yang dibutuhkan oleh aplikasi Anda.
|
|
Evaluation configuration |
|
|
Preset template evaluation |
Tersedia beberapa templat preset. Memilih beberapa templat akan mengagregasi hasil di halaman detail tugas. Contoh berikut menggunakan Answer Correctness Evaluation:
Parameter utama:
|
|
Custom evaluation |
Buat evaluasi kustom dengan templat prompt yang ditentukan pengguna.
|
|
Resource configuration: Digunakan untuk penjadwalan tugas evaluasi. Pilih CPU resources berdasarkan kompleksitas tugas. |
|
Lihat hasil evaluasi
Setelah mengirimkan tugas evaluasi, LangStudio akan mengarahkan Anda ke halaman Overview tugas tersebut. Setiap eksekusi terdiri dari dua tahap: Batch Run memproses setiap baris dataset melalui aplikasi Anda, dan Metric Evaluation memberi skor pada setiap output terhadap bidang tambahan. Setelah selesai, Anda dapat melihat jejak, metrik, dan detail output untuk setiap subtugas.

Halaman Metrics menampilkan semua hasil metrik evaluasi. Nama metrik ditentukan dalam Lampiran: Templat evaluasi preset.

Lampiran: Templat evaluasi preset
LangStudio menyediakan templat evaluasi bawaan berikut:
|
Nama templat |
Deskripsi |
Jenis layanan model |
Bidang input |
|
Exact Match Evaluation |
Membandingkan output Agent dengan referensi untuk kecocokan eksak. Skor: 0 (tidak cocok) hingga 1 (cocok sempurna). |
None |
|
|
Answer Relevancy Evaluation |
Memberi skor relevansi output aplikasi terhadap input menggunakan juri LLM. Skor: 1–5 (lebih tinggi = lebih relevan). |
LLM |
|
|
Answer Correctness Evaluation |
Menilai akurasi faktual, cakupan informasi, dan kecocokan format terhadap referensi. Skor: 1–5 (lebih tinggi = lebih mendekati kecocokan). |
LLM |
|
|
Instruction Following Evaluation |
Menilai kepatuhan terhadap instruksi dalam konten, format, dan batasan. Skor: 1–5 (lebih tinggi = kepatuhan lebih baik). |
LLM |
|
|
Answer Faithfulness Evaluation |
Mendeteksi informasi yang dibuat-buat yang tidak didukung atau bertentangan dengan konteks. Skor: 1–5 (lebih tinggi = lebih setia). |
LLM |
|
|
Safety Evaluation |
Mendeteksi konten berbahaya, ofensif, atau tidak pantas. Skor: 1–5 (lebih tinggi = lebih aman). |
LLM |
|
|
Trajectory Evaluation |
Menilai lintasan eksekusi Agent secara holistik. Skor: 1–5 (lebih tinggi = kinerja lebih baik). |
LLM |
|


