予測フォーマット
予測にはJSON形式で表現されるOpenLabelフォーマットを使用します。これはプレアノテーションプレアノテーションのアップロードに使用されるフォーマットと同じです。OpenLabelフォーマットの一般的な情報については、OpenLABEL形式OpenLABEL形式をご覧ください。
サポートされる予測機能
現在の予測アップロードAPIは、以下のジオメトリをサポートしています:
名前 | OpenLABELフィールド | 説明 |
|---|---|---|
キューボイド | cuboid | 3Dキューボイド |
バウンディングボックス | bbox | 2Dバウンディングボックス |
ビットマップ(セグメンテーション) | image | 画像用セグメンテーションビットマップ |
キューボイドの回転は、エクスポート時(OpenLABEL形式OpenLABEL形式)と同じにしてください(詳細は座標系座標系を参照)。2Dジオメトリはピクセル座標で表現してください。
このAPIでは、関連する部分(キー)はframes、objects、streams、ontologies、およびmetadataです。最後のmetadataは最も簡単で、schema_version": "1.0.0"と記述するだけです(完全なコンテキストについては以下の例を参照)。また、streamも簡単で、どのセンサー(カメラ、LiDARなど)があるか、およびその名前を指定します(例:sensor_name: {"type": "camera"}やsensor_name: {"type": "lidar"})。こちらも完全なコンテキストについては以下の例を参照してください。
シーケンス全体で時間的に変化する予測のすべての部分(座標や動的プロパティなど)はframesに記述されます。シーケンスの各フレームはframes内のキーと値のペアで表現されます。キーはframe_idで、値は以下のようになります:
frame_id: {
"frame_properties": {
"timestamp": 0,
"external_id": "",
"streams": {}
},
"objects": {
...
}
}frame_properties.timestampの値(ミリ秒単位、非シーケンスデータの場合は0に設定することを推奨)は、各予測フレームを該当するアノテーション済みフレームとマッチングするために使用されるため、アノテーションされたシーンと一致する必要があります。frame_id(文字列)は、基となるシーンの記述に使用されるframe_idに従うことを推奨しますが、不一致の場合はframe_properties.timestampが優先されます。非シーケンスデータの場合、frame_idには"0"が適切です。frame_properties.external_idとframe_properties.streamの値は、図のように空のままにすると自動的に解決されます。
objectsキーには、キーと値のペアが含まれ、各ペアは基本的にそのフレーム内の1つのオブジェクトを表します。objectsキーは各フレーム内とルートの両方に存在することに注意してください。基本的に同じオブジェクトを記述しますが、時間的に変化する可能性のある情報(座標など、フレーム固有の情報)はフレームに属し、静的な情報(オブジェクトクラスなど)はルートに属します。オブジェクトキー(文字列)は任意ですが、同じオブジェクトを記述する場合は、異なるobjects内のキーが一致する必要があります。
オブジェクトの詳細な記述方法については、以下の例を参照してください。キューボイドとバウンディングボックスには、フレーム固有の属性confidenceを指定することで存在信頼度を提供できます。値は0.0から1.0の間の数値である必要があり、空の場合は1.0に設定されます。指定する場合は、数値として定義する必要があります。静的なobject_data.typeは、ツール内でクラス名として表示されます。
セグメンテーションビットマップの場合、画像自体はアノテーション対象画像と同じ解像度のグレースケール8ビットPNG画像です(実際の予測がアノテーション画像を部分的にしかカバーしていない場合や解像度が低い場合は、パディングやアップスケーリングが必要です)。画像自体は、base64エンコードされた文字列としてフレームのオブジェクトにペーストすることでOpenLabel内に提供されます。以下の例を参照してください。さらに、各色レベルに対応するクラスを記述するontologyも提供する必要があります。8ビットグレースケール画像では、最大256クラスをエンコードできます。セグメンテーション以外の予測では、ontologyは省略できます。
以下の例のcamera_idは、アノテーション済みシーンのセンサーIDと一致する必要があります。一方、LiDARセンサーの対応するIDは@lidarに設定してください。
予測の例
静的プロパティcolorを持つ2フレームの2Dバウンディングボックス
OpenLabelでは、バウンディングボックスは4つの値のリスト[x, y, width, height]として表現されます。xとyはバウンディングボックスの中心座標です。widthとheightはバウンディングボックスの幅と高さです。xとyの座標は画像の左上隅を基準とします。
{
"openlabel": {
"frames": {
"0": {
"frame_properties": {
"timestamp": 0,
"external_id": "",
"streams": {}
},
"objects": {
"1232b4f4-e3ca-446a-91cb-d8d403703df7": {
"object_data": {
"bbox": [
{
"attributes": {
"num": [
{
"val": 0.85,
"name": "confidence"
}
],
"text": [
{
"name": "stream",
"val": "camera_id"
}
]
},
"name": "any-human-readable-bounding-box-name",
"val": [
1.0,
1.0,
40.0,
30.0
]
}
]
}
}
}
},
"1": {
"frame_properties": {
"timestamp": 50,
"external_id": "",
"streams": {}
},
"objects": {
"1232b4f4-e3ca-446a-91cb-d8d403703df7": {
"object_data": {
"bbox": [
{
"attributes": {
"num": [
{
"val": 0.82,
"name": "confidence"
}
],
"text": [
{
"name": "stream",
"val": "camera_id"
}
]
},
"name": "any-human-readable-bounding-box-name",
"val": [
2.0,
3.0,
30.0,
20.0
]
}
]
}
}
}
}
},
"metadata": {
"schema_version": "1.0.0"
},
"objects": {
"1232b4f4-e3ca-446a-91cb-d8d403703df7": {
"name": "any-human-readable-bounding-box-name",
"object_data": {
"text": [
{
"name": "color",
"val": "red"
}
]
},
"type": "PassengerCar"
}
},
"streams": {
"camera_id": {
"type": "camera"
}
}
}
}静的プロパティcolorを持つ2フレームの3Dキューボイド
キューボイドは10個の値のリスト[x, y, z, qx, qy, qz, qw, width, length, height]として表現されます。x、y、zはキューボイドの中心座標です。x、y、z、width、length、heightの単位はメートルです。qx、qy、qz、qwはキューボイドの回転を表すクォータニオン値です。
座標系とクォータニオンの詳細については、OpenLABEL形式OpenLABEL形式をご覧ください。
{
"openlabel": {
"frames": {
"0": {
"frame_properties": {
"timestamp": 0,
"external_id": "",
"streams": {}
},
"objects": {
"1232b4f4-e3ca-446a-91cb-d8d403703df7": {
"object_data": {
"cuboid": [
{
"attributes": {
"num": [
{
"val": 0.85,
"name": "confidence"
}
],
"text": [
{
"name": "stream",
"val": "@lidar"
}
]
},
"name": "any-human-readable-cuboid-name",
"val": [
2.079312801361084,
-18.919870376586914,
0.3359137773513794,
-0.002808041640852679,
0.022641949116037438,
0.06772797660868829,
0.9974429197838155,
1.767102435869269,
4.099334155319101,
1.3691029802958168
]
}
]
}
}
}
},
"1": {
"frame_properties": {
"timestamp": 50,
"external_id": "",
"streams": {}
},
"objects": {
"1232b4f4-e3ca-446a-91cb-d8d403703df7": {
"object_data": {
"cuboid": [
{
"attributes": {
"num": [
{
"val": 0.87,
"name": "confidence"
}
],
"text": [
{
"name": "stream",
"val": "@lidar"
}
]
},
"name": "any-human-readable-cuboid-name",
"val": [
3.123312801361927,
-20.285740376586913,
0.0649137773513349,
-0.002808041640852679,
0.022641949116037438,
0.06772797660868829,
0.9974429197838155,
1.767102435869269,
4.099334155319101,
1.3691029802958168
]
}
]
}
}
}
}
},
"metadata": {
"schema_version": "1.0.0"
},
"objects": {
"1232b4f4-e3ca-446a-91cb-d8d403703df7": {
"name": "any-human-readable-cuboid-name",
"object_data": {
"text": [
{
"name": "color",
"val": "red"
}
]
},
"type": "PassengerCar"
}
},
"streams": {
"@lidar": {
"type": "lidar"
}
}
}
}単一フレームのセグメンテーションビットマップ
Python PILを使用した小さなカラー画像の変換、アップスケーリング、パディング、およびbase64エンコードによる大きなグレースケール画像への変換
このコード例は、解像度300 x 200のマルチカラー予測ビットマップ画像を、解像度1000 x 800のグレースケール画像に変換する方法を示しています。まずグレースケールに変換し、次に予測を600 x 400にリスケーリングし、両側に均等にパディングします。また、OpenLabelで使用できるように、画像をbase64エンコードされた文字列に変換するコードも含まれています。このコードは組み込みのnumpy関数のみを使用しており、パフォーマンスの最適化は行われていません。
import base64
import io
import numpy as np
from PIL import Image
# The original mapping used to produce the images
original_mapping = {
(0,0,0): "_background",
(255,0,0): "class_1",
(0,0,255): "class_2",
}
# The grayscale mapping (this will also be the ontology in the openlabel)
grayscale_mapping = {
"_background": 0,
"class_1": 1,
"class_2": 2,
}
prediction = Image.open("my_original_prediction_file.png") # Let's say this has resolution 300 x 200
def lookup(pixel_color):
return grayscale_mapping[original_mapping[tuple(pixel_color)]]
# convert to grayscale via numpy array lookup
prediciton_numpy = np.array(prediction)
grayscale_prediction_numpy = np.vectorize(lookup, signature="(m)->()")(prediciton_numpy)
grayscale_prediction = Image.fromarray(grayscale_prediction_numpy.astype(np.uint8))
# upscale to another resolution
upscaled_grayscale_prediction = grayscale_prediction.resize((600, 400), resample=Image.Resampling.NEAREST)
# padding by first constructing a new background image of target size, and then paste the prediction in the right position
padded_grayscale_prediction = Image.new("L", (1000, 800), 0)
padded_grayscale_prediction.paste(upscaled_grayscale_prediction, (201, 201))
image_bytes = io.BytesIO()
padded_grayscale_prediction.save(image_bytes, format="PNG")
prediction_str = base64.b64encode(image_bytes.getvalue()).decode("utf-8")
セグメンテーションビットマップ用のOpenLabel
prediction_strとgrayscale_mappingは、その後OpenLabel内で以下のように使用できます:
{
"openlabel": {
"frames": {
"0": {
"objects": {
"07d469f9-c9ab-44ec-8d09-0c72bdb44dc2": {
"object_data": {
"image": [
{
"name": "a_human_readable_name",
"val": prediction_str,
"mime_type": "image/png",
"encoding": "base64",
"attributes": {
"text": [
{
"val": "camera_id",
"name": "stream"
}
]
}
}
]
}
}
},
"frame_properties": {
"streams": {},
"timestamp": 0,
"external_id": ""
},
}
},
"objects": {
"07d469f9-c9ab-44ec-8d09-0c72bdb44dc2": {
"name": "07d469f9-c9ab-44ec-8d09-0c72bdb44dc2",
"type": "segmentation_bitmap"
}
},
"streams": {
"camera_id": {
"type": "camera"
}
},
"metadata": {
"schema_version": "1.0.0"
},
"ontologies": {
"0": {
"classifications": {str(v): k for k, v in grayscale_mapping.items()},
"uri": ""
}
}
}
}シーン内の複数のカメラに対して予測を提供する場合は、画像のリストを拡張できます。
kognic-openlabelを使用してフォーマットを検証
詳細についてはkognic-openlabelを参照してください。