All Products
Search
Document Center

Platform For AI:Evaluasi aplikasi

Last Updated:Aug 26, 2026

Evaluasi kinerja alur aplikasi Anda sebelum diterapkan di lingkungan produksi. LangStudio memberi skor aplikasi berdasarkan beberapa dimensi menggunakan templat evaluasi preset atau kustom.

Ikhtisar

Untuk mengevaluasi alur aplikasi, konfigurasikan set data evaluasi, petakan bidang input, dan pilih templat evaluasi. Sistem memproses setiap baris dataset secara batch melalui aplikasi Anda, memberi skor pada setiap output terhadap bidang tambahan, lalu mengagregasi hasilnya.

image

Sebelum memulai

  • Buat dan debug alur aplikasi. Kembangkan alur aplikasi.

  • Siapkan set data evaluasi dalam format JSONL dan unggah ke OSS. Contoh:

    {"history":[],"query": "Describe the perilous majesty of Mount Hua", "reference": "Mount Hua stands alone, soaring to the clouds; \nSheer cliffs cut the sky, with rugged, handsome crags. \nGreen pines and bamboo vie for beauty on the cliffs; \nMonkeys cry and eagles fly, lit by frosty swords of light. \n\nPerilous peaks like scissors, jagged swords pointing to the sky; \nNarrow paths on steep slopes, where vines are the only way. \nWind and mist intertwine, as clouds emerge from caves; \nA deep fairyland, with a heavenly ladder hard to climb. \n\nJagged ridges cross, like a surging dragon's spine; \nDangerous paths lead onward, twisting toward the heavens. \nFrom lonely pine tops, eagles strike the vast sky; \nAt the summit of Mount Hua, a majestic and heroic sight.", "contexts": ["Mount Hua is one of the Five Great Mountains of China.", "Mount Hua is famous for its precipitous cliffs."]}
    {"history":[],"query": "Can you list 5 rare metals? Please rank them by global demand.", "reference": "Rare metals are metallic elements that are scarce in the Earth's crust, unevenly distributed, or difficult to mine. They play a crucial role in high-tech fields and emerging industries. The ranking of global demand can change with time and technological progress, but the following are some rare metals that are typically in high demand. This list is not necessarily ranked by absolute demand, as that can vary at different times.\n\n1. **Cobalt (Co)** - Cobalt is a key component of lithium-ion batteries, especially in electric vehicles and portable electronics. It is also used to manufacture heat-resistant alloys, hard alloys, and catalysts.\n\n2. **Neodymium (Nd)** - Neodymium is a rare-earth metal mainly used to produce strong magnets, such as high-performance permanent magnets. These magnets are widely used in computer hard drives, wind turbines, and the drive motors of electric vehicles.\n\n3. **Lithium (Li)** - Lithium is primarily used to manufacture lithium batteries. As the demand for electric vehicles and portable electronic devices increases, the demand for lithium is rising rapidly.\n\n4. **Silver (Ag)** - Although silver is not as rare as the metals listed above, its industrial demand is huge. It is mainly used in electronics, solar panels, jewelry, and currency manufacturing.\n\n5. **Ruthenium (Ru)** - Ruthenium is a rare precious metal widely used for data storage in hard disk drives and large-capacity servers. It is also used in catalysts and electrochemical cells.\n\nThe demand for these metals is influenced by many factors, such as the global economy, technological development, and policy support. Moreover, as time passes and markets change, other rare metals such as tantalum, indium, rhenium, and other rare-earth metals may also appear on the list of most in-demand rare metals.", "contexts": ["Rare metals are metals with low abundance in the Earth's crust that are complex to mine and extract.", "Lithium (Li): Used in battery manufacturing.", "Cobalt (Co): Used in high-performance alloys and battery manufacturing."]}

    File contoh: langstudio_eval_demo.jsonl

  • Buat koneksi LLM yang diperlukan untuk evaluasi. Konfigurasi koneksi.

    Catatan: Beberapa templat evaluasi memerlukan model LLM-as-a-Judge. Konfigurasikan koneksi LLM yang sesuai sebelum menggunakan templat tersebut.

Penagihan

Evaluasi alur aplikasi menggunakan OSS untuk penyimpanan dataset dan PAI-DLC untuk menjalankan tugas offline. Anda dikenai biaya untuk kedua resource tersebut. Penagihan OSS. Penagihan Deep Learning Containers (DLC).

Buat tugas evaluasi

Setelah men-debug aplikasi Anda, klik Evaluation di pojok kanan atas untuk membuat tugas evaluasi.

image

Parameter utama:

Parameter

Deskripsi

Evaluation dataset

OSS file

Pilih file JSONL dari OSS sebagai set data evaluasi Anda. Dataset harus berisi bidang question sebagai input aplikasi, ditambah bidang lain yang dibutuhkan oleh templat evaluasi yang Anda pilih. Lampiran: Templat evaluasi preset.

Application flow input mapping

question/chat_history

Petakan bidang input aplikasi ke kolom dataset.

Catatan: Tugas evaluasi menjalankan aplikasi Anda untuk inferensi sebelum memberi skor hasil. Pilih bidang input yang dibutuhkan oleh aplikasi Anda.

  • Dalam mode workflow, bidang input ditentukan oleh node awal. Bidang input biasanya berupa question dan chat_history.

  • Dalam mode kode, satu-satunya bidang input adalah question.

Evaluation configuration

Preset template evaluation

Tersedia beberapa templat preset. Memilih beberapa templat akan mengagregasi hasil di halaman detail tugas. Contoh berikut menggunakan Answer Correctness Evaluation:

image

Parameter utama:

  • model_configuration: Model LLM-as-a-Judge yang memberi skor kesesuaian antara question dan answer dari aplikasi. Pilih model yang mumpuni seperti qwen3-max.

  • reference: Kolom referensi dari set data evaluasi. Templat Answer Correctness Evaluation menggunakan question, answer, dan reference untuk menghasilkan skor kebenaran. LangStudio menangkap question dan answer secara otomatis — Anda hanya perlu menentukan kolom referensi. Catatan: Sebagian besar templat hanya memerlukan model LLM-as-a-Judge; bidang reference bersifat opsional.

Lampiran: Templat evaluasi preset.

Custom evaluation

Buat evaluasi kustom dengan templat prompt yang ditentukan pengguna.

image

image

  • Anda dapat membuat beberapa evaluasi kustom. Masing-masing harus memiliki nama unik.

  • Konfigurasikan templat evaluasi dengan judge_template. Contoh disediakan. Aturan:

    • Jangan ubah deskripsi format output dalam prompt.

    • Tiga placeholder bawaan: {query} merepresentasikan kueri, {messages} merepresentasikan proses eksekusi Agent, dan {response} merepresentasikan respons Agent.

    • Gunakan {data.***} untuk mereferensikan bidang dataset. Misalnya, {data.judge} mereferensikan bidang judge. Lihat contoh Dynamic Rule Evaluation.

    • Karakter {} merupakan karakter khusus yang merepresentasikan variabel. Jika Anda perlu menyertakan karakter ini dalam prompt, gunakan {{}}.

Resource configuration: Digunakan untuk penjadwalan tugas evaluasi. Pilih CPU resources berdasarkan kompleksitas tugas.

Lihat hasil evaluasi

Setelah mengirimkan tugas evaluasi, LangStudio akan mengarahkan Anda ke halaman Overview tugas tersebut. Setiap eksekusi terdiri dari dua tahap: Batch Run memproses setiap baris dataset melalui aplikasi Anda, dan Metric Evaluation memberi skor pada setiap output terhadap bidang tambahan. Setelah selesai, Anda dapat melihat jejak, metrik, dan detail output untuk setiap subtugas.

image

Halaman Metrics menampilkan semua hasil metrik evaluasi. Nama metrik ditentukan dalam Lampiran: Templat evaluasi preset.

image

Lampiran: Templat evaluasi preset

LangStudio menyediakan templat evaluasi bawaan berikut:

Nama templat

Deskripsi

Jenis layanan model

Bidang input

Exact Match Evaluation

Membandingkan output Agent dengan referensi untuk kecocokan eksak. Skor: 0 (tidak cocok) hingga 1 (cocok sempurna).

None

  • reference: Referensi. Jenis data: String.

Answer Relevancy Evaluation

Memberi skor relevansi output aplikasi terhadap input menggunakan juri LLM. Skor: 1–5 (lebih tinggi = lebih relevan).

LLM

  • model_configuration: Model LLM-as-a-Judge.

Answer Correctness Evaluation

Menilai akurasi faktual, cakupan informasi, dan kecocokan format terhadap referensi. Skor: 1–5 (lebih tinggi = lebih mendekati kecocokan).

LLM

  • model_configuration: Model LLM-as-a-Judge.

  • reference (Opsional): Referensi. Jenis data: String.

Instruction Following Evaluation

Menilai kepatuhan terhadap instruksi dalam konten, format, dan batasan. Skor: 1–5 (lebih tinggi = kepatuhan lebih baik).

LLM

  • model_configuration: Model LLM-as-a-Judge.

Answer Faithfulness Evaluation

Mendeteksi informasi yang dibuat-buat yang tidak didukung atau bertentangan dengan konteks. Skor: 1–5 (lebih tinggi = lebih setia).

LLM

  • model_configuration: Model LLM-as-a-Judge.

  • contexts: Konteks. Jenis data: List[String].

Safety Evaluation

Mendeteksi konten berbahaya, ofensif, atau tidak pantas. Skor: 1–5 (lebih tinggi = lebih aman).

LLM

  • model_configuration: Model LLM-as-a-Judge.

Trajectory Evaluation

Menilai lintasan eksekusi Agent secara holistik. Skor: 1–5 (lebih tinggi = kinerja lebih baik).

LLM

  • model_configuration: Model LLM-as-a-Judge.