Future-State Predicting LSTM for Early Surgery Type Recognition

by   Siddharth Kannan, et al.

This work presents a novel approach for the early recognition of the type of a laparoscopic surgery from its video. Early recognition algorithms can be beneficial to the development of 'smart' OR systems that can provide automatic context-aware assistance, and also enable quick database indexing. The task is however ridden with challenges specific to videos belonging to the domain of laparoscopy, such as high visual similarity across surgeries and large variations in video durations. To capture the spatio-temporal dependencies in these videos, we choose as our model a combination of a Convolutional Neural Network (CNN) and Long Short-Term Memory (LSTM) network. We then propose two complementary approaches for improving early recognition performance. The first approach is a CNN fine-tuning method that encourages surgeries to be distinguished based on the initial frames of laparoscopic videos. The second approach, referred to as 'Future-State Predicting LSTM', trains an LSTM to predict information related to future frames, which helps in distinguishing between the different types of surgeries. We evaluate our approaches on a large dataset of 425 laparoscopic videos containing 9 types of surgeries (Laparo425), and achieve on average an accuracy of 75 minutes of a surgery. These results are quite promising from a practical standpoint and also encouraging for other types of image-guided surgeries.


page 1

page 5

page 9


Predicting Video Saliency with Object-to-Motion CNN and Two-layer Convolutional LSTM

Over the past few years, deep neural networks (DNNs) have exhibited grea...

Proof of Concept: Automatic Type Recognition

The type used to print an early modern book can give scholars valuable i...

Predicting tongue motion in unlabeled ultrasound videos using convolutional LSTM neural network

A challenge in speech production research is to predict future tongue mo...

Recognizing and Curating Photo Albums via Event-Specific Image Importance

Automatic organization of personal photos is a problem with many real wo...

Player Identification in Hockey Broadcast Videos

We present a deep recurrent convolutional neural network (CNN) approach ...

Please sign up or login with your details

Forgot password? Click here to reset