예측 형식
예측(prediction)은 JSON으로 표현되는 OpenLabel 포맷을 사용합니다. 이는 사전 어노테이션(pre-annotation)을 업로드할 때 사용하는 것과 동일한 포맷입니다. OpenLabel 포맷에 대한 일반적인 정보는 여기에서 확인할 수 있습니다.
지원되는 예측 기능
현재 예측 업로드 API는 다음 형상(geometry)을 지원합니다.
Name | OpenLABEL field | Description |
|---|---|---|
Cuboid | cuboid | 3D 큐보이드 |
Bounding box | bbox | 2D 바운딩 박스 |
Bitmaps (segmentation) | image | 이미지에 대한 세그멘테이션 비트맵 |
이 API에서 관련된 부분(키)은 frames, objects, streams, ontologies, metadata입니다. 마지막 항목인 metadata가 가장 간단하며, 그냥 schema_version": "1.0.0"으로 작성하면 됩니다(전체 맥락은 아래 예제 참고). stream 역시 간단하며, 카메라, 라이다 등 어떤 센서들이 있는지와 그 이름을 명시해야 합니다. 예를 들어 sensor_name: {"type": "camera"} 또는 sensor_name: {"type": "lidar"}와 같습니다. 전체 맥락은 아래 예제를 참고하십시오.
시퀀스 전체에 걸쳐 시간에 따라 변하는 예측의 모든 부분, 예를 들어 좌표나 동적 속성 등은 frames에서 기술됩니다. 시퀀스의 각 프레임은 frames 아래의 키-값 쌍으로 표현됩니다. 키는 frame_id이며, 값은 다음과 같은 형태여야 합니다.
frame_id: {
"frame_properties": {
"timestamp": 0,
"external_id": "",
"streams": {}
},
"objects": {
...
}
}frame_properties.timestamp 값(ms 단위이며, 시퀀스가 아닌 데이터의 경우 0으로 설정하는 것을 권장)은 각 예측된 프레임을 관련 어노테이션 프레임과 매칭하는 데 사용되므로, 어노테이션된 씬과 일치해야 합니다. frame_id(문자열)는 기본 씬을 기술하는 데 사용된 frame_id를 따르는 것을 권장하지만, 불일치가 있는 경우 frame_properties.timestamp가 우선합니다. 시퀀스가 아닌 데이터의 경우 frame_id로 "0"을 사용하는 것이 좋습니다. frame_properties.external_id와 frame_properties.stream 값은 예시처럼 비워 두면 자동으로 결정됩니다.
objects 키는 다시 키-값 쌍을 포함하며, 각 쌍은 기본적으로 해당 프레임의 객체 하나를 나타냅니다. 각 프레임뿐만 아니라 루트에도 objects 키가 존재한다는 점에 유의하십시오. 두 곳은 기본적으로 동일한 객체들을 설명하지만, 좌표처럼 시간에 따라 변할 수 있는(즉 프레임별) 정보는 프레임에 속하고, 객체 클래스와 같은 정적인 정보는 루트에 속합니다. 객체 키(문자열)는 임의로 정할 수 있지만, 동일한 객체를 설명하는 경우 서로 다른 objects에서 키가 일치해야 합니다.
객체를 상세히 기술하는 방법은 아래 예제를 참고하십시오. 큐보이드와 바운딩 박스의 경우, 프레임별 속성인 confidence를 지정하여 존재 신뢰도(existence confidence)를 제공할 수 있습니다. 이 값은 0.0에서 1.0 사이의 숫자 값이어야 하며, 비워 두면 1.0으로 설정됩니다. 값을 제공하는 경우 반드시 숫자 값으로 정의해야 합니다. 정적 필드인 object_data.type은 도구에서 클래스 이름으로 표시됩니다.
세그멘테이션 비트맵의 경우, 이미지 자체는 어노테이션된 이미지와 동일한 해상도의 그레이스케일 8비트 PNG 이미지여야 합니다(실제 예측이 어노테이션된 이미지의 일부만 커버하거나 해상도가 더 낮은 경우, 패딩 및/또는 업스케일링을 해야 합니다). 이미지 자체는 openlabel 내에서 프레임에 속한 객체에 base64 인코딩 문자열로 붙여넣는 방식으로 제공됩니다. 아래 예제를 참고하십시오. 또한 각 색상 레벨이 어떤 클래스에 대응하는지를 설명하는 ontology도 제공해야 합니다. 8비트 그레이스케일 이미지에서는 최대 256개의 클래스를 인코딩할 수 있습니다. 세그멘테이션이 아닌 예측의 경우 ontology는 생략할 수 있습니다.
아래 예제의 camera_id는 어노테이션된 씬 내 센서의 id와 일치해야 하며, 라이다 센서에 해당하는 id는 @lidar로 설정해야 합니다.
예측 예제
정적 속성 color를 가진 두 프레임의 2D 바운딩 박스
OpenLabel에서 바운딩 박스는 [x, y, width, height]라는 4개의 값을 가진 리스트로 표현됩니다. 여기서 x와 y는 바운딩 박스 중심 좌표입니다. width와 height는 바운딩 박스의 너비와 높이입니다. x와 y 좌표는 이미지의 왼쪽 상단 모서리를 기준으로 합니다.
{
"openlabel": {
"frames": {
"0": {
"frame_properties": {
"timestamp": 0,
"external_id": "",
"streams": {}
},
"objects": {
"1232b4f4-e3ca-446a-91cb-d8d403703df7": {
"object_data": {
"bbox": [
{
"attributes": {
"num": [
{
"val": 0.85,
"name": "confidence"
}
],
"text": [
{
"name": "stream",
"val": "camera_id"
}
]
},
"name": "any-human-readable-bounding-box-name",
"val": [
1.0,
1.0,
40.0,
30.0
]
}
]
}
}
}
},
"1": {
"frame_properties": {
"timestamp": 50,
"external_id": "",
"streams": {}
},
"objects": {
"1232b4f4-e3ca-446a-91cb-d8d403703df7": {
"object_data": {
"bbox": [
{
"attributes": {
"num": [
{
"val": 0.82,
"name": "confidence"
}
],
"text": [
{
"name": "stream",
"val": "camera_id"
}
]
},
"name": "any-human-readable-bounding-box-name",
"val": [
2.0,
3.0,
30.0,
20.0
]
}
]
}
}
}
}
},
"metadata": {
"schema_version": "1.0.0"
},
"objects": {
"1232b4f4-e3ca-446a-91cb-d8d403703df7": {
"name": "any-human-readable-bounding-box-name",
"object_data": {
"text": [
{
"name": "color",
"val": "red"
}
]
},
"type": "PassengerCar"
}
},
"streams": {
"camera_id": {
"type": "camera"
}
}
}
}정적 속성 color를 가진 두 프레임의 3D 큐보이드
큐보이드는 [x, y, z, qx, qy, qz, qw, width, length, height]라는 10개의 값을 가진 리스트로 표현됩니다. 여기서 x, y, z는 큐보이드 중심 좌표입니다. x, y, z, width, length, height는 미터 단위입니다. qx, qy, qz, qw는 큐보이드 회전을 나타내는 쿼터니언 값입니다.
좌표계와 쿼터니언에 대해 더 알아보려면 여기를 참고하십시오.
{
"openlabel": {
"frames": {
"0": {
"frame_properties": {
"timestamp": 0,
"external_id": "",
"streams": {}
},
"objects": {
"1232b4f4-e3ca-446a-91cb-d8d403703df7": {
"object_data": {
"cuboid": [
{
"attributes": {
"num": [
{
"val": 0.85,
"name": "confidence"
}
],
"text": [
{
"name": "stream",
"val": "@lidar"
}
]
},
"name": "any-human-readable-cuboid-name",
"val": [
2.079312801361084,
-18.919870376586914,
0.3359137773513794,
-0.002808041640852679,
0.022641949116037438,
0.06772797660868829,
0.9974429197838155,
1.767102435869269,
4.099334155319101,
1.3691029802958168
]
}
]
}
}
}
},
"1": {
"frame_properties": {
"timestamp": 50,
"external_id": "",
"streams": {}
},
"objects": {
"1232b4f4-e3ca-446a-91cb-d8d403703df7": {
"object_data": {
"cuboid": [
{
"attributes": {
"num": [
{
"val": 0.87,
"name": "confidence"
}
],
"text": [
{
"name": "stream",
"val": "@lidar"
}
]
},
"name": "any-human-readable-cuboid-name",
"val": [
3.123312801361927,
-20.285740376586913,
0.0649137773513349,
-0.002808041640852679,
0.022641949116037438,
0.06772797660868829,
0.9974429197838155,
1.767102435869269,
4.099334155319101,
1.3691029802958168
]
}
]
}
}
}
}
},
"metadata": {
"schema_version": "1.0.0"
},
"objects": {
"1232b4f4-e3ca-446a-91cb-d8d403703df7": {
"name": "any-human-readable-cuboid-name",
"object_data": {
"text": [
{
"name": "color",
"val": "red"
}
]
},
"type": "PassengerCar"
}
},
"streams": {
"@lidar": {
"type": "lidar"
}
}
}
}단일 프레임 세그멘테이션 비트맵
Python PIL을 사용하여 작은 컬러 이미지를 변환, 업스케일링, 패딩 및 base64 인코딩하여 큰 그레이스케일 이미지로 만들기
이 코드 예제는 해상도 300 x 200의 다색 예측 비트맵 이미지를 해상도 1000 x 800의 그레이스케일 이미지로 변환하는 방법을 보여줍니다. 먼저 그레이스케일로 변환한 다음, 예측을 600 x 400으로 리스케일하고, 양쪽에 균등하게 패딩을 추가합니다. 또한 이미지를 base64 문자열로 인코딩하는 코드도 포함되어 있으며, 이는 이후 openlabel에서 사용할 수 있습니다. 이 코드는 numpy의 내장 함수만 사용하며 성능에 최적화되어 있지는 않습니다.
import base64
import io
import numpy as np
from PIL import Image
# The original mapping used to produce the images
original_mapping = {
(0,0,0): "_background",
(255,0,0): "class_1",
(0,0,255): "class_2",
}
# The grayscale mapping (this will also be the ontology in the openlabel)
grayscale_mapping = {
"_background": 0,
"class_1": 1,
"class_2": 2,
}
prediction = Image.open("my_original_prediction_file.png") # Let's say this has resolution 300 x 200
def lookup(pixel_color):
return grayscale_mapping[original_mapping[tuple(pixel_color)]]
# convert to grayscale via numpy array lookup
prediciton_numpy = np.array(prediction)
grayscale_prediction_numpy = np.vectorize(lookup, signature="(m)->()")(prediciton_numpy)
grayscale_prediction = Image.fromarray(grayscale_prediction_numpy.astype(np.uint8))
# upscale to another resolution
upscaled_grayscale_prediction = grayscale_prediction.resize((600, 400), resample=Image.Resampling.NEAREST)
# padding by first constructing a new background image of target size, and then paste the prediction in the right position
padded_grayscale_prediction = Image.new("L", (1000, 800), 0)
padded_grayscale_prediction.paste(upscaled_grayscale_prediction, (201, 201))
image_bytes = io.BytesIO()
padded_grayscale_prediction.save(image_bytes, format="PNG")
prediction_str = base64.b64encode(image_bytes.getvalue()).decode("utf-8")
세그멘테이션 비트맵을 위한 Openlabel
prediction_str과 grayscale_mapping은 이후 다음과 같이 openlabel에서 사용할 수 있습니다.
{
"openlabel": {
"frames": {
"0": {
"objects": {
"07d469f9-c9ab-44ec-8d09-0c72bdb44dc2": {
"object_data": {
"image": [
{
"name": "a_human_readable_name",
"val": prediction_str,
"mime_type": "image/png",
"encoding": "base64",
"attributes": {
"text": [
{
"val": "camera_id",
"name": "stream"
}
]
}
}
]
}
}
},
"frame_properties": {
"streams": {},
"timestamp": 0,
"external_id": ""
},
}
},
"objects": {
"07d469f9-c9ab-44ec-8d09-0c72bdb44dc2": {
"name": "07d469f9-c9ab-44ec-8d09-0c72bdb44dc2",
"type": "segmentation_bitmap"
}
},
"streams": {
"camera_id": {
"type": "camera"
}
},
"metadata": {
"schema_version": "1.0.0"
},
"ontologies": {
"0": {
"classifications": {str(v): k for k, v in grayscale_mapping.items()},
"uri": ""
}
}
}
}씬 내 여러 카메라에 대한 예측을 제공하는 경우, 이미지 목록을 확장하면 됩니다.
kognic-openlabel을 사용하여 포맷 검증하기
자세한 내용은 kognic-openlabel을 참고하십시오.