本番デプロイの前に、アプリケーションフローのパフォーマンスを評価します。LangStudio は、プリセットまたはカスタムの評価テンプレートを使用して、複数の側面からアプリケーションをスコアリングします。
概要
アプリケーションフローを評価するには、評価データセットを設定し、入力フィールドをマッピングし、評価テンプレートを選択します。システムは、データセットの各行をアプリケーション経由でバッチ処理し、補助フィールドを基に各出力をスコアリングし、結果を集計します。

事前準備
-
アプリケーションフローを作成し、デバッグします。詳細については、「アプリケーションフローの開発」をご参照ください。
-
JSONL 形式で評価データセットを準備し、OSS にアップロードします。例:
{"history":[],"query": "Describe the perilous majesty of Mount Hua", "reference": "Mount Hua stands alone, soaring to the clouds; \nSheer cliffs cut the sky, with rugged, handsome crags. \nGreen pines and bamboo vie for beauty on the cliffs; \nMonkeys cry and eagles fly, lit by frosty swords of light. \n\nPerilous peaks like scissors, jagged swords pointing to the sky; \nNarrow paths on steep slopes, where vines are the only way. \nWind and mist intertwine, as clouds emerge from caves; \nA deep fairyland, with a heavenly ladder hard to climb. \n\nJagged ridges cross, like a surging dragon's spine; \nDangerous paths lead onward, twisting toward the heavens. \nFrom lonely pine tops, eagles strike the vast sky; \nAt the summit of Mount Hua, a majestic and heroic sight.", "contexts": ["Mount Hua is one of the Five Great Mountains of China.", "Mount Hua is famous for its precipitous cliffs."]} {"history":[],"query": "Can you list 5 rare metals? Please rank them by global demand.", "reference": "Rare metals are metallic elements that are scarce in the Earth's crust, unevenly distributed, or difficult to mine. They play a crucial role in high-tech fields and emerging industries. The ranking of global demand can change with time and technological progress, but the following are some rare metals that are typically in high demand. This list is not necessarily ranked by absolute demand, as that can vary at different times.\n\n1. **Cobalt (Co)** - Cobalt is a key component of lithium-ion batteries, especially in electric vehicles and portable electronics. It is also used to manufacture heat-resistant alloys, hard alloys, and catalysts.\n\n2. **Neodymium (Nd)** - Neodymium is a rare-earth metal mainly used to produce strong magnets, such as high-performance permanent magnets. These magnets are widely used in computer hard drives, wind turbines, and the drive motors of electric vehicles.\n\n3. **Lithium (Li)** - Lithium is primarily used to manufacture lithium batteries. As the demand for electric vehicles and portable electronic devices increases, the demand for lithium is rising rapidly.\n\n4. **Silver (Ag)** - Although silver is not as rare as the metals listed above, its industrial demand is huge. It is mainly used in electronics, solar panels, jewelry, and currency manufacturing.\n\n5. **Ruthenium (Ru)** - Ruthenium is a rare precious metal widely used for data storage in hard disk drives and large-capacity servers. It is also used in catalysts and electrochemical cells.\n\nThe demand for these metals is influenced by many factors, such as the global economy, technological development, and policy support. Moreover, as time passes and markets change, other rare metals such as tantalum, indium, rhenium, and other rare-earth metals may also appear on the list of most in-demand rare metals.", "contexts": ["Rare metals are metals with low abundance in the Earth's crust that are complex to mine and extract.", "Lithium (Li): Used in battery manufacturing.", "Cobalt (Co): Used in high-performance alloys and battery manufacturing."]}サンプルファイル:langstudio_eval_demo.jsonl
-
評価に必要な LLM 接続を作成します。詳細については、「接続設定」をご参照ください。
注意:一部の評価テンプレートでは、「ジャッジとしての LLM」モデルが必要です。これらのテンプレートを使用する前に、対応する LLM 接続を設定してください。
課金
アプリケーションフロー評価では、データセットのストレージには OSS を、オフラインタスクの実行には PAI-DLC を使用します。これらの両方のリソースに課金されます。詳細については、「OSS の課金」と「Deep Learning Containers (DLC) の課金」をご参照ください。
評価タスクの作成
アプリケーションをデバッグした後、右上隅にある 評価 をクリックして評価タスクを作成します。

主なパラメーター:
|
パラメーター |
説明 |
|
評価データセット |
|
|
OSS ファイル |
OSS から JSONL ファイルを選択し、評価データセットとして使用します。データセットには、アプリケーションの入力として |
|
アプリケーションフロー入力マッピング |
|
|
question/chat_history |
アプリケーションの入力フィールドをデータセットの列にマッピングします。 注意:評価タスクは、結果をスコアリングする前に、推論のためにアプリケーションを実行します。アプリケーションで必要な入力フィールドを選択してください。
|
|
評価設定 |
|
|
プリセットテンプレート評価 |
複数のプリセットテンプレートを利用できます。複数のテンプレートを選択すると、タスク詳細ページで結果が集計されます。次の例では、[回答正確性評価] を使用しています:
主なパラメーター:
詳細については、「付録:プリセット評価テンプレート」をご参照ください。 |
|
カスタム評価 |
ユーザー定義のプロンプトテンプレートを使用して、カスタム評価を作成します。
|
|
リソース設定:評価タスクのスケジューリングに使用されます。タスクの複雑さに応じて CPU リソース を選択してください。 |
|
評価結果の表示
評価タスクを送信すると、LangStudio はタスクの 概要 ページにリダイレクトします。各実行には 2 つのステージがあります。 [Batch Run] はデータセットの各行をアプリケーション経由で処理し、 [Metric Evaluation] は補助フィールドを基に各出力をスコアリングします。完了後、各サブタスクのトレース、メトリック、および出力詳細を表示できるようになります。

メトリクス ページには、すべての評価メトリックの結果が表示されます。メトリック名は、付録:プリセット評価テンプレートで定義されています。

付録:プリセット評価テンプレート
LangStudio は、次の組み込み評価テンプレートを用意しています:
|
テンプレート名 |
説明 |
モデルサービスタイプ |
入力フィールド |
|
完全一致評価 |
Agent の出力を参照と比較し、完全一致を評価します。スコア:0 (不一致) から 1 (完全一致)。 |
なし |
|
|
回答関連性評価 |
LLM ジャッジを使用して、アプリケーションの出力と入力との関連性をスコアリングします。スコア:1~5 (高いほど関連性が高い)。 |
LLM |
|
|
回答正確性評価 |
参照と比較して、事実の正確性、情報の網羅性、および形式の一致度を評価します。スコア:1~5 (高いほど一致度が高い)。 |
LLM |
|
|
指示追従性評価 |
内容、形式、制約における指示への準拠度を評価します。スコア:1~5 (高いほど準拠度が高い)。 |
LLM |
|
|
回答忠実性評価 |
コンテキストによって裏付けられていない、または矛盾する捏造された情報を検出します。スコア:1~5 (高いほど忠実性が高い)。 |
LLM |
|
|
安全性評価 |
有害、攻撃的、または不適切なコンテンツを検出します。スコア:1~5 (高いほど安全性が高い)。 |
LLM |
|
|
トラジェクトリ評価 |
Agent の実行トラジェクトリを総合的に評価します。スコア:1~5 (高いほどパフォーマンスが高い)。 |
LLM |
|


