ConceptioArchiveGoogle Patents
Google Patentsopen access

Event/object-of-interest centric timelapse video generation on camera device … — Ambarella International Lp (US11594254B2)

Ambarella International Lp · Google Patents
Google Patents · Patents · License: Open Access
Open Source ↗
ambarellainternationallp
patent, google patents, intellectual property, US11594254B2, Ambarella International Lp, Luyi Sun, en, 2023

ABSTRACT

Abstract

An apparatus including an interface and a processor. The interface may be configured to receive pixel data generated by a capture device. The processor may be configured to generate video frames in response to the pixel data, perform computer vision operations on the video frames to detect objects, perform a classification of the objects detected based on characteristics of the objects, determine whether the classification of the objects corresponds to a user-defined event and generate encoded video frames from the video frames. The encoded video frames may be communicated to a cloud storage service. The encoded video frames may comprise a first sample of the video frames selected at a first rate when the user-defined event is not detected and a second sample of the video frames selected at a second rate while the user-defined event is detected. The second rate may be greater than the first rate.

Description

This application relates to U.S. patent application Ser. No. 17/126,108, filed on Dec. 18, 2020, which relates to China Patent Application No. 202010836914.7, filed on Aug. 19, 2020. Each of the mentioned applications are hereby incorporated by reference in its entirety.

FIELD OF THE INVENTION

The invention relates to computer vision generally and, more particularly, to a method and/or apparatus for implementing an event/object-of-interest centric timelapse video generation on camera device with the assistance of neural network input.

BACKGROUND

A timelapse video mode for conventional internet-connected/cloud-enabled cameras is usually implemented by software operated on a cloud server. The software relies on using the resources of distributed processing of the cloud server (i.e., scalable computing). Timelapse video clips are displayed at a fixed frame rate for a fast-forward effect. For example, video frames are selected at fixed intervals to create the timelapse video (i.e., every thirtieth frame is selected from a thirty frames per second video to create the timelapse video).

Internet-connected/cloud-enabled cameras encode video data and then communicate the encoded video streams to the cloud servers. In order to transcode the encoded video streams into a timelapse video, the cloud server has to first decode the encoded video stream. The cloud servers spend lots of CPU cycles to decode the conventional video streams (i.e., compressed video using AVC or HEVC encoding), extract the video frames at fixed frame intervals, then transcode these frames into a timelapse video. Meanwhile, even with a timelapse video, users have difficulty finding important details captured by the internet-connected/cloud-enabled cameras. Since a timelapse video always uses a fixed frame rate that does not use all the video data originally captured by the internet-connected/cloud-enabled, video frames at a normal display speed are not available for the entire duration of time that something of interest to the user is in the original video captured. For example, when a security camera observes a potential event of interest, such as a person on the premises, the timelapse will be the same as when the security camera observes nothing of particular interest.

It would be desirable to implement an event/object-of-interest centric timelapse video generation on camera device with the assistance of neural network input.

SUMMARY

The invention concerns an apparatus comprising an interface and a processor. The interface may be configured to receive pixel data generated by a capture device. The processor may be configured to receive the pixel data from the interface, generate video frames in response to the pixel data, perform computer vision operations on the video frames to detect objects, perform a classification of the objects detected based on characteristics of the objects, determine whether the classification of the objects corresponds to a user-defined event and generate encoded video frames from the video frames. The encoded video frames may be communicated to a cloud storage service. The encoded video frames may comprise a first sample of the video frames selected at a first rate when the user-defined event is not detected and a second sample of the video frames selected at a second rate while the user-defined event is detected. The second rate may be greater than the first rate.

BRIEF DESCRIPTION OF THE FIGURES

Embodiments of the invention will be apparent from the following detailed description and the appended claims and drawings.

FIG. 1 is a diagram illustrating an example context of the present invention.

FIG. 2 is a diagram illustrating example internet-connected cameras implementing an example embodiment of the present invention.

FIG. 3 is a block diagram illustrating components of an apparatus configured to provide an event centric timelapse video with the assistance of a neural network.

FIG. 4 is a diagram illustrating an interconnected camera communicating with a cloud server and a video processing pipeline for generating a timelapse video.

FIG. 5 is a diagram illustrating a smart timelapse mode on an edge AI camera with CV analysis using all frames of an event.

FIG. 6 is a diagram illustrating a smart timelapse mode on an edge AI camera with CV analysis using partial frames of an event.

FIG. 7 is a diagram illustrating event detection in a captured video frame.

FIG. 8 is a diagram illustrating an application operating on a smart phone for controlling preferences for a timelapse video.

FIG. 9 is a flow diagram illustrating a method for implementing an event/object-of-interest centric timelapse video generation on a camera device with the assistance of neural network input.

FIG. 10 is a flow diagram illustrating a method for selecting a video frames in response to a determination of whether an event has been detected.

FIG. 11 is a flow diagram illustrating a method for updating a feature set for event detection in response to user input.

FIG. 12 is a flow diagram illustrating a method for implementing an event/object-of-interest centric timelapse video generation using scalable computing resources.

DETAILED DESCRIPTION OF THE EMBODIMENTS

Embodiments of the present invention include providing an event/object-of-interest centric timelapse video generation on camera device with the assistance of neural network input that may (i) implement event detection on an edge device, (ii) implement video encoding on an edge device, (iii) generate smart timelapse videos with varying framerates based on events/objects detected, (iv) detect objects/events using a convolutional neural network implemented on a processor, (v) adjust a timelapse framerate to capture all video frames when an event/object is detected, (vi) perform facial recognition and/or object classification local to the camera device, (vii) upload an encoded smart timelapse video to a cloud storage server, (viii) enable configuration of parameters for event/object detection and/or a framerate of a timelapse video using a smartphone app and/or (ix) be implemented as one or more integrated circuits.

Embodiments of the present invention may be configured to generate a smart timelapse video. The smart timelapse video may be implemented by adjusting a video display speed automatically based on content detected in the captured video. The video display speed may be adjusted in response to a selection rate of video frames from the originally captured video to be used for the smart timelapse video. Embodiments of the present invention may be configured to generate encoded smart timelapse videos. The smart timelapse video may be generated on a cloud service (e.g., using scalable computing). The smart timelapse video may be generated on an edge device (e.g., an artificial intelligence (AI) camera). Generating the smart timelapse video on the edge device may comprise only using processing capabilities within the edge device (e.g., without outsourcing processing to an external device). In an example, the edge device may be an internet-connected/cloud-enabled camera.

The edge AI device/camera may be configured to implement artificial intelligence (AI) technology. Using AI technology, the edge AI camera may be a more powerful (e.g., by providing relevant data for the user) and a more power efficient solution than using a cloud server in many aspects. An edge AI camera may be configured to execute computer readable instructions locally (e.g., internally) on the device (e.g., without relying on external processing resources) to analyze the video content frame by frame. Based on the analysis, content in the video frames may be tagged with metadata information. The metadata information may be used to select video frames for a smart timelapse video. For example, the video frames may be classified by being tagged as having no interesting object/event and as having an interesting object/event. Computer vision (CV) operations may determine whether there is an interesting object/event (e.g., based on a pre-defined feature set).

When there is no interesting CV event (or type of object) in the video frames for a duration of N seconds (N=60/120/190/ . . . ), the edge AI camera may select one of the video frames captured during the no-event duration. The selected no-event video frame may be used for video encoding (e.g., a video encoder built in the edge AI camera device that performs video encoding internally). Selecting one video frame from the no-event duration N for encoding may be repeated for each N second duration with no event detected. Selecting one video frame from the no-event duration N for encoding may result in an encoded output that provides a condensed portion of the captured video (e.g., to effectively fast-forward through “meaningless content” portions of video captured with a high display speed).

When there is an interesting CV event (or type of object) detected in the video frames for a duration of M seconds (M=5/15/30/ . . . ), the edge AI camera may adjust how many and/or the rate of selection of the CV event video frames for the event duration M. The event detected may be defined based on a pre-determined feature set. In an example, the object and/or event of interest may be considered to be detected when a person is detected, a car is detected, an animal detected, an amount of motion is detected, a particular face detected, etc. The selection of the video frames for the smart timelapse video may be adjusted to select all of the video frames from the M second duration of the CV event (e.g., for a 2 minute event captured at 60 frames per second, the entire 7200 frames may be selected). The selection of the CV event video frames may be adjusted to select video frames at a higher rate than for the no-event duration (e.g., select more frames, but not all frames) for the M second duration of the event (e.g., for a 2 minute event captured at 60 frames per second, the rate of selection may be changed to 30 frames and every other frame may be selecting resulting in 3600 video frames being selected).

The selected CV event video frames may be encoded (e.g., using the on-device video encoding of the AI edge camera to perform encoding locally). Selecting video frames at a higher rate for the event duration M for encoding may result in an encoded output that provides either a portion of the smart timelapse video with the normal display speed for the “meaningful content” or a portion of the smart timelapse video with a slightly condensed (but not as high speed as the “meaningless content”) display speed for the “meaningful content”.

With the smart timelapse video generation implemented on an edge camera, the user may quickly browse through the captured video content for a long period of time (e.g., days/weeks/months), and the user may be confident that no interesting CV event will be missed. An interesting CV event may comprise detecting a person (e.g., a known person) by face detection and/or face recognition, detecting a car license plate (e.g., using a license plate reader), detecting a pet (e.g., a known animal) using animal/pet recognition, detecting motion (e.g., any motion detected above a pre-defined threshold), etc. The type of event and/or object detected that may be considered an interesting event may be varied according to the design criteria of a particular implementation.

Embodiments of the present invention may enable a user to specify the type of object/event that may be considered to be interesting. In one example, an app operating on a smartphone (e.g., a companion app for the edge AI camera) may be configured to adjust settings for the edge AI camera. In another example, the edge AI camera may be configured to provide a web interface (e.g., over a local area network) to enable a user to remotely select objects/events to be considered interesting events. In yet another example, the edge AI camera may connect to a cloud server, and the user may use a web interface to adjust settings stored on the cloud server that may then be sent to the edge AI camera to control the smart timelapse object type(s) of interest.

The duration of the smart timelapse video may be configured as regular intervals and/or since the last time the user has received a timelapse video. For example, if the user has missed 20 event notifications, the moment the users interacts with the app (e.g., swipes on a smartphone) to view the events, the event-focused timelapse may be presented to the user for viewing. While embodiments of the present invention may generally be performed local to an edge AI camera (e.g., implementing a processor configured to implement a convolutional neural network and/or video encoding locally), embodiments of the present invention may be performed using software on the cloud server to achieve similar effects.

Referring to FIG. 1 , a diagram illustrating an example context of the present invention is shown. A home 50 and a vehicle 52 are shown. Camera systems 100 a - 100 n are shown. Each of the cameras 100 a - 100 n may be configured to generate pixel data of the environment, generate video frames from the pixel data, encode the video frames and/or generate the smart timelapse videos. For example, each of the cameras 100 a - 100 n may be configured to operate independently of each other. Each of the cameras 100 a - 100 n may capture video and generate smart timelapse videos. In one example, the respective smart timelapse videos may be uploaded to a cloud storage service. In another example, the respective smart timelapse videos may be stored locally (e.g., on a microSD card, to a local network attached storage device, etc.).

Each of the cameras 100 a - 100 n may be configured to detect different or the same events/objects that may be considered interesting. For example, the camera system 100 b may capture an area near an entrance of the home 50 . For an entrance of the home 50 , objects/events of interest may be detecting people. The camera system 100 b may be configured to analyze video frames to detect people and the smart timelapse video may slow down (e.g., select video frames for encoding at a higher frame rate) when a person is detected. In another example, the camera system 100 d may capture an area near the vehicle</figure-call

This application relates to U.S. patent application Ser. No. 17/126,108, filed on Dec. 18, 2020, which relates to China Patent Application No. 202010836914.7, filed on Aug. 19, 2020. Each of the mentioned applications are hereby incorporated by reference in its entirety.

FIELD OF THE INVENTION

The invention relates to computer vision generally and, more particularly, to a method and/or apparatus for implementing an event/object-of-interest centric timelapse video generation on camera device with the assistance of neural network input.

BACKGROUND

A timelapse video mode for conventional internet-connected/cloud-enabled cameras is usually implemented by software operated on a cloud server. The software relies on using the resources of distributed processing of the cloud server (i.e., scalable computing). Timelapse video clips are displayed at a fixed frame rate for a fast-forward effect. For example, video frames are selected at fixed intervals to create the timelapse video (i.e., every thirtieth frame is selected from a thirty frames per second video to create the timelapse video).

Internet-connected/cloud-enabled cameras encode video data and then communicate the encoded video streams to the cloud servers. In order to transcode the encoded video streams into a timelapse video, the cloud server has to first decode the encoded video stream. The cloud servers spend lots of CPU cycles to decode the conventional video streams (i.e., compressed video using AVC or HEVC encoding), extract the video frames at fixed frame intervals, then transcode these frames into a timelapse video. Meanwhile, even with a timelapse video, users have difficulty finding important details captured by the internet-connected/cloud-enabled cameras. Since a timelapse video always uses a fixed frame rate that does not use all the video data originally captured by the internet-connected/cloud-enabled, video frames at a normal display speed are not available for the entire duration of time that something of interest to the user is in the original video captured. For example, when a security camera observes a potential event of interest, such as a person on the premises, the timelapse will be the same as when the security camera observes nothing of particular interest.

It would be desirable to implement an event/object-of-interest centric timelapse video generation on camera device with the assistance of neural network input.

SUMMARY

The invention concerns an apparatus comprising an interface and a processor. The interface may be configured to receive pixel data generated by a capture device. The processor may be configured to receive the pixel data from the interface, generate video frames in response to the pixel data, perform computer vision operations on the video frames to detect objects, perform a classification of the objects detected based on characteristics of the objects, determine whether the classification of the objects corresponds to a user-defined event and generate encoded video frames from the video frames. The encoded video frames may be communicated to a cloud storage service. The encoded video frames may comprise a first sample of the video frames selected at a first rate when the user-defined event is not detected and a second sample of the video frames selected at a second rate while the user-defined event is detected. The second rate may be greater than the first rate.

BRIEF DESCRIPTION OF THE FIGURES

Embodiments of the invention will be apparent from the following detailed description and the appended claims and drawings.

FIG. 1 is a diagram illustrating an example context of the present invention.

FIG. 2 is a diagram illustrating example internet-connected cameras implementing an example embodiment of the present invention.

FIG. 3 is a block diagram illustrating components of an apparatus configured to provide an event centric timelapse video with the assistance of a neural network.

FIG. 4 is a diagram illustrating an interconnected camera communicating with a cloud server and a video processing pipeline for generating a timelapse video.

FIG. 5 is a diagram illustrating a smart timelapse mode on an edge AI camera with CV analysis using all frames of an event.

FIG. 6 is a diagram illustrating a smart timelapse mode on an edge AI camera with CV analysis using partial frames of an event.

FIG. 7 is a diagram illustrating event detection in a captured video frame.

FIG. 8 is a diagram illustrating an application operating on a smart phone for controlling preferences for a timelapse video.

FIG. 9 is a flow diagram illustrating a method for implementing an event/object-of-interest centric timelapse video generation on a camera device with the assistance of neural network input.

FIG. 10 is a flow diagram illustrating a method for selecting a video frames in response to a determination of whether an event has been detected.

FIG. 11 is a flow diagram illustrating a method for updating a feature set for event detection in response to user input.

FIG. 12 is a flow diagram illustrating a method for implementing an event/object-of-interest centric timelapse video generation using scalable computing resources.

DETAILED DESCRIPTION OF THE EMBODIMENTS

Embodiments of the present invention include providing an event/object-of-interest centric timelapse video generation on camera device with the assistance of neural network input that may (i) implement event detection on an edge device, (ii) implement video encoding on an edge device, (iii) generate smart timelapse videos with varying framerates based on events/objects detected, (iv) detect objects/events using a convolutional neural network implemented on a processor, (v) adjust a timelapse framerate to capture all video frames when an event/object is detected, (vi) perform facial recognition and/or object classification local to the camera device, (vii) upload an encoded smart timelapse video to a cloud storage server, (viii) enable configuration of parameters for event/object detection and/or a framerate of a timelapse video using a smartphone app and/or (ix) be implemented as one or more integrated circuits.

Embodiments of the present invention may be configured to generate a smart timelapse video. The smart timelapse video may be implemented by adjusting a video display speed automatically based on content detected in the captured video. The video display speed may be adjusted in response to a selection rate of video frames from the originally captured video to be used for the smart timelapse video. Embodiments of the present invention may be configured to generate encoded smart timelapse videos. The smart timelapse video may be generated on a cloud service (e.g., using scalable computing). The smart timelapse video may be generated on an edge device (e.g., an artificial intelligence (AI) camera). Generating the smart timelapse video on the edge device may comprise only using processing capabilities within the edge device (e.g., without outsourcing processing to an external device). In an example, the edge device may be an internet-connected/cloud-enabled camera.

The edge AI device/camera may be configured to implement artificial intelligence (AI) technology. Using AI technology, the edge AI camera may be a more powerful (e.g., by providing relevant data for the user) and a more power efficient solution than using a cloud server in many aspects. An edge AI camera may be configured to execute computer readable instructions locally (e.g., internally) on the device (e.g., without relying on external processing resources) to analyze the video content frame by frame. Based on the analysis, content in the video frames may be tagged with metadata information. The metadata information may be used to select video frames for a smart timelapse video. For example, the video frames may be classified by being tagged as having no interesting object/event and as having an interesting object/event. Computer vision (CV) operations may determine whether there is an interesting object/event (e.g., based on a pre-defined feature set).

When there is no interesting CV event (or type of object) in the video frames for a duration of N seconds (N=60/120/190/ . . . ), the edge AI camera may select one of the video frames captured during the no-event duration. The selected no-event video frame may be used for video encoding (e.g., a video encoder built in the edge AI camera device that performs video encoding internally). Selecting one video frame from the no-event duration N for encoding may be repeated for each N second duration with no event detected. Selecting one video frame from the no-event duration N for encoding may result in an encoded output that provides a condensed portion of the captured video (e.g., to effectively fast-forward through “meaningless content” portions of video captured with a high display speed).

When there is an interesting CV event (or type of object) detected in the video frames for a duration of M seconds (M=5/15/30/ . . . ), the edge AI camera may adjust how many and/or the rate of selection of the CV event video frames for the event duration M. The event detected may be defined based on a pre-determined feature set. In an example, the object and/or event of interest may be considered to be detected when a person is detected, a car is detected, an animal detected, an amount of motion is detected, a particular face detected, etc. The selection of the video frames for the smart timelapse video may be adjusted to select all of the video frames from the M second duration of the CV event (e.g., for a 2 minute event captured at 60 frames per second, the entire 7200 frames may be selected). The selection of the CV event video frames may be adjusted to select video frames at a higher rate than for the no-event duration (e.g., select more frames, but not all frames) for the M second duration of the event (e.g., for a 2 minute event captured at 60 frames per second, the rate of selection may be changed to 30 frames and every other frame may be selecting resulting in 3600 video frames being selected).

The selected CV event video frames may be encoded (e.g., using the on-device video encoding of the AI edge camera to perform encoding locally). Selecting video frames at a higher rate for the event duration M for encoding may result in an encoded output that provides either a portion of the smart timelapse video with the normal display speed for the “meaningful content” or a portion of the smart timelapse video with a slightly condensed (but not as high speed as the “meaningless content”) display speed for the “meaningful content”.

With the smart timelapse video generation implemented on an edge camera, the user may quickly browse through the captured video content for a long period of time (e.g., days/weeks/months), and the user may be confident that no interesting CV event will be missed. An interesting CV event may comprise detecting a person (e.g., a known person) by face detection and/or face recognition, detecting a car license plate (e.g., using a license plate reader), detecting a pet (e.g., a known animal) using animal/pet recognition, detecting motion (e.g., any motion detected above a pre-defined threshold), etc. The type of event and/or object detected that may be considered an interesting event may be varied according to the design criteria of a particular implementation.

Embodiments of the present invention may enable a user to specify the type of object/event that may be considered to be interesting. In one example, an app operating on a smartphone (e.g., a companion app for the edge AI camera) may be configured to adjust settings for the edge AI camera. In another example, the edge AI camera may be configured to provide a web interface (e.g., over a local area network) to enable a user to remotely select objects/events to be considered interesting events. In yet another example, the edge AI camera may connect to a cloud server, and the user may use a web interface to adjust settings stored on the cloud server that may then be sent to the edge AI camera to control the smart timelapse object type(s) of interest.

The duration of the smart timelapse video may be configured as regular intervals and/or since the last time the user has received a timelapse video. For example, if the user has missed 20 event notifications, the moment the users interacts with the app (e.g., swipes on a smartphone) to view the events, the event-focused timelapse may be presented to the user for viewing. While embodiments of the present invention may generally be performed local to an edge AI camera (e.g., implementing a processor configured to implement a convolutional neural network and/or video encoding locally), embodiments of the present invention may be performed using software on the cloud server to achieve similar effects.

Referring to FIG. 1 , a diagram illustrating an example context of the present invention is shown. A home 50 and a vehicle 52 are shown. Camera systems 100 a - 100 n are shown. Each of the cameras 100 a - 100 n may be configured to generate pixel data of the environment, generate video frames from the pixel data, encode the video frames and/or generate the smart timelapse videos. For example, each of the cameras 100 a - 100 n may be configured to operate independently of each other. Each of the cameras 100 a - 100 n may capture video and generate smart timelapse videos. In one example, the respective smart timelapse videos may be uploaded to a cloud storage service. In another example, the respective smart timelapse videos may be stored locally (e.g., on a microSD card, to a local network attached storage device, etc.).

Each of the cameras 100 a - 100 n may be configured to detect different or the same events/objects that may be considered interesting. For example, the camera system 100 b may capture an area near an entrance of the home 50 . For an entrance of the home 50 , objects/events of interest may be detecting people. The camera system 100 b may be configured to analyze video frames to detect people and the smart timelapse video may slow down (e.g., select video frames for encoding at a higher frame rate) when a person is detected. In another example, the camera system 100 d may capture an area near the vehicle 52 . For the vehicle 52 , objects/events of interest may be detecting other vehicles and pedestrians. The camera system 100 b may be configured to analyze video frames to detect vehicles (or road signs) and people and the smart timelapse video may slow down when a vehicle or a pedestrian is detected.

Each of the cameras 100 a - 100 n may operate independently from each other. For example, each of the cameras 100 a - 100 n may individually analyze the pixel data captured and perform the event/object detection locally. In some embodiments, the cameras 100 a - 100 n may be configured as a network of cameras (e.g., security cameras that send video data to a central source such as network-attached storage and/or a cloud service). The locations and/or configurations of the cameras 100 a - 100 n may be varied according to the design criteria of a particular implementation.

Referring to FIG. 2 , a diagram illustrating example internet-connected cameras implementing an example embodiment of the present invention is shown. Camera systems 100 a - 100 n are shown. Each camera device 100 a - 100 n may have a different style and/or use case. For example, the camera 100 a may be an action camera, the camera 100 b may be a ceiling mounted security camera, the camera 100 n may be webcam, etc. Other types of cameras may be implemented (e.g., home security cameras, battery powered cameras, doorbell cameras, stereo cameras, etc.). The design/style of the cameras 100 a - 100 n may be varied according to the design criteria of a particular implementation.

Each of the camera systems 100 a - 100 n may comprise a block (or circuit) 102 and/or a block (or circuit) 104 . The circuit 102 may implement a processor. The circuit 104 may implement a capture device. The camera systems 100 a - 100 n may comprise other components (not shown). Details of the components of the cameras 100 a - 100 n may be described in association with FIG. 3 .

The processor 102 may be configured to implement a convolutional neural network (CNN). The processor 102 may be configured to implement a video encoder. The processor 102 may generate the smart timelapse videos. The capture device 104 may be configured to capture pixel data that may be used by the processor 102 to generate video frames.

The cameras 100 a - 100 n may be edge devices. The processor 102 implemented by each of the cameras 100 a - 100 n may enable the cameras 100 a - 100 n to implement various functionality internally (e.g., at a local level). For example, the processor 102 may be configured to perform object/event detection (e.g., computer vision operations), video encoding and/or video transcoding on-device. For example, even advanced processes such as computer vision may be performed by the processor 102 without uploading video data to a cloud service in order to offload computation-heavy functions (e.g., computer vision, video encoding, video transcoding, etc.).

Referring to FIG. 3 , a block diagram illustrating components of an apparatus configured to provide an event centric timelapse video with the assistance of a neural network is shown. A block diagram of the camera system 100 i is shown. The camera system 100 i may be a representative example of the camera system 100 a - 100 n shown in association with FIGS. 1 - 2 . The camera system 100 i generally comprises the processor 102 , the capture devices 104 a - 104 n , blocks (or circuits) 150 a - 150 n , a block (or circuit) 152 , blocks (or circuits) 154 a - 154 n , a block (or circuit) 156 , blocks (or circuits) 158 a - 158 n , a block (or circuit) 160 and/or a block (or circuit) 162 . The blocks 150 a - 150 n may implement lenses. The circuit 152 may implement sensors. The circuits 154 a - 154 n may implement microphones (e.g., audio capture devices). The circuit 156 may implement a communication device. The circuits 158 a - 158 n may implement audio output devices (e.g., speakers). The circuit 160 may implement a memory. The circuit 162 may implement a power supply (e.g., a battery). The camera system 100 i may comprise other components (not shown). In the example shown, some of the components 150 - 158 are shown external to the camera system 100 i . However, the components 150 - 158 may be implemented within and/or attached to the camera system 100 i (e.g., the speakers 158 a - 158 n may provide better functionality if not located inside a housing of the camera system 100 i ). The number, type and/or arrangement of the components of the camera system 100 i may be varied according to the design criteria of a particular implementation.

In an example implementation, the processor 102 may be implemented as a video processor. The processor 102 may comprise inputs 170 a - 170 n and/or other inputs. The processor 102 may comprise an input/ output 172 . The processor 102 may comprise an input 174 and an input 176 . The processor 102 may comprise an output 178 . The processor 102 may comprise an output 180 a and an input 180 b . The number of inputs, outputs and/or bi-directional ports implemented by the processor 102 may be varied according to the design criteria of a particular implementation.

In the embodiment shown, the capture devices 104 a - 104 n may be components of the camera system 100 i . In some embodiments, the capture devices 104 a - 104 n may be separate devices (e.g., remotely connected to the camera system 100 i , such as a drone, a robot and/or a system of security cameras configured capture video data) configured to send data to the camera system 100 i . In one example, the capture devices 104 a - 104 n may be implemented as part of an autonomous robot configured to patrol particular paths such as hallways. Similarly, in the example shown, the sensors 152 , the microphones 154 a - 154 n , the wireless communication device 156 , and/or the speakers 158 a - 158 n are shown external to the camera system 100 i but in some embodiments may be a component of (e.g., within) the camera system 100 i.

The camera system 100 i may receive one or more signals (e.g., IMF_A-IMF_N), a signal (e.g., SEN), a signal (e.g., FEAT_SET) and/or one or more signals (e.g., DIR_AUD). The camera system 100 i may present a signal (e.g., ENC_VIDEO) and/or a signal (e.g., DIR_AOUT). The capture devices 104 a - 104 n may receive the signals IMF_A-IMF_N from the corresponding lenses 150 a - 150 n . The processor 102 may receive the signal SEN from the sensors 152 . The processor 102 may receive the signal DIR_AUD from the microphones 154 a - 154 n . The processor 102 may present the signal ENC_VIDEO to the communication device 156 and receive the signal FEAT_SET from the communication device 156 . For example, the wireless communication device 156 may be a radio-frequency (RF) transmitter. In another example, the communication device 156 may be a Wi-Fi module. In another example, the communication device 156 may be a device capable of implementing RF transmission, Wi-Fi, Bluetooth and/or other wireless communication protocols. In some embodiments, the signal ENC_VIDEO may be presented to a display device connected to the camera 100 i . The processor 102 may present the signal DIR_AOUT to the speakers 158 a - 158 n.

The lenses 150 a - 150 n may capture signals (e.g., IM_A-IM_N). The signals IM_A-IM_N may be an image (e.g., an analog image) of the environment near the camera system 100 i presented by the lenses 150 a - 150 n to the capture devices 104 a - 104 n as the signals IMF_A-IMF_N. The lenses 150 a - 150 n may be implemented as an optical lens. The lenses 150 a - 150 n may provide a zooming feature and/or a focusing feature. The capture devices 104 a - 104 n and/or the lenses 150 a - 150 n may be implemented, in one example, as a single lens assembly. In another example, the lenses 150 a - 150 n may be a separate implementation from the capture devices 104 a - 104 n . The capture devices 104 a - 104 n are shown within the circuit 100 i . In an example implementation, the capture devices 104 a - 104 n may be implemented outside of the circuit 100 i (e.g., along with the lenses 150 a - 150 n as part of a lens/capture device assembly).

In some embodiments, two or more of the lenses 150 a - 150 n may be configured as a stereo pair of lenses. For example, the camera 100 i may implement stereo vision. The lenses 150 a - 150 n implemented as a stereo pair may be implemented at a pre-determined distance apart from each other and at a pre-determined inward angle. The pre-determined distance and/or the pre-determined inward angle may be used by the processor 102 to build disparity maps for stereo vision.

The capture devices 104 a - 104 n may be configured to capture image data for video (e.g., the signals IMF_A-IMF_N from the lenses 150 a - 150 n ). In some embodiments, the capture devices 104 a - 104 n may be video capturing devices such as cameras. The capture devices 104 a - 104 n may capture data received through the lenses 150 a - 150 n to generate raw pixel data. In some embodiments, the capture devices 104 a - 104 n may capture data received through the lenses 150 a - 150 n to generate bitstreams (e.g., generate video frames). For example, the capture devices 104 a - 104 n may receive focused light from the lenses 150 a - 150 n . The lenses 150 a - 150 n may be directed, tilted, panned, zoomed and/or rotated to provide a targeted view from the camera system 100 i (e.g., a view for a video frame, a view for a panoramic video frame captured using multiple capture devices 104 a - 104 n , a target image and reference image view for stereo vision, etc.). The capture devices 104 a - 104 n may generate signals (e.g., PIXELD_A-PIXELD_N). The signals PIXELD_A-PIXELD_N may be pixel data (e.g., a sequence of pixels that may be used to generate video frames). In some embodiments, the signals PIXELD_A-PIXELD_N may be video data (e.g., a sequence of video frames). The signals PIXELD_A-PIXELD_N may be presented to the inputs 170 a - 170 n of the processor 102 .

The capture devices 104 a - 104 n may transform the received focused light signals IMF_A-IMF_N into digital data (e.g., bitstreams). In some embodiments, the capture devices 104 a - 104 n may perform an analog to digital conversion. For example, the capture devices 104 a - 104 n may perform a photoelectric conversion of the focused light received by the lenses 150 a - 150 n . The capture devices 104 a - 104 n may transform the bitstreams into pixel data, images and/or video frames. In some embodiments, the pixel data generated by the capture devices 104 a - 104 n may be uncompressed and/or raw data generated in response to the focused light from the lenses 150 a - 150 n . In some embodiments, the output of the capture devices 104 a - 104 n may be digital video signals.

The sensors 152 may comprise one or more input devices. The sensors 152 may be configured to detect physical input from the environment and convert the physical input into computer readable signals. The signal SEN may comprise the computer readable signals generated by the sensors 152 . In an example, one of the sensors 152 may be configured to detect an amount of light and present a computer readable signal representing the amount of light detected. In another example, one of the sensors 152 may be configured to detect motion and present a computer readable signal representing the amount of motion detected. The sensors 152 may be configured to detect temperature (e.g., a thermometer), orientation (e.g., a gyroscope), a movement speed (e.g., an accelerometer), etc. The types of input detected by the sensors 152 may be varied according to the design criteria of a particular implementation.

The data provided in the signal SEN provided by the sensors 152 may be read and/or interpreted by the processor 102 . The processor 102 may use the data provided by the signal SEN for various operations. In some embodiments, the processor 102 may use a light reading from the sensors 152 to determine whether to activate an infrared light (e.g., to provide night vision). In another example, the processor 102 may use information about movement from an accelerometer and/or a gyroscope to perform motion correction on video frames generated. The types of operations performed by the processor 102 in response to the signal SEN may be varied according to the design criteria of a particular implementation.

The communication device 156 may send and/or receive data to/from the camera system 100 i . In some embodiments, the communication device 156 may be implemented as a wireless communications module. In some embodiments, the communication device 156 may be implemented as a satellite connection to a proprietary system. In one example, the communication device 156 may be a hard-wired data port (e.g., a USB port, a mini-USB port, a USB-C connector, HDMI port, an Ethernet port, a DisplayPort interface, a Lightning port, etc.). In another example, the communication device 156 may be a wireless data interface (e.g., Wi-Fi, Bluetooth, ZigBee, cellular, etc.).

The communication device 156 may be configured to receive the signal FEAT_SET. The signal FEAT_SET may comprise a feature set. The feature set received may be used to detect events and/or objects. For example, the feature set may be used to perform the computer vision operations. The feature set information may comprise instructions for the processor 102 for determining which types of objects correspond to an object and/or event of interest.

The processor 102 may receive the signals PIXELD_A-PIXELD_N from the capture devices 104 a - 104 n at the inputs 170 a - 170 n . The processor 102 may send/receive a signal (e.g., DATA) to/from the memory 160 at the input/ output 172 . The processor 102 may receive the signal SEN from the sensors 152 at the input port 174 . The processor 102 may receive the signal DIR_AUD from the microphones 154 a - 154 n at the port 176 . The processor 102 may send the signal DIR_AOUT to the speakers 158 a - 158 n via the port 178 . The processor 102 may send the signal ENC_VIDEO to the communication device 156 via the output port 180 a . The processor 102 may receive the signal FEAT_SET from the communication device 156 via the input port 180 b . In an example, the processor 102 may be connected through a bi-directional interface (or connection) to the capture devices 104 a - 104 n , the sensors 152 , the microphones 154 a - 154 n , the communication device 156 , and/or the speakers 158 a - 158 n and/or the memory 160 . The processor 102 may store and/or retrieve data from the memory 160 . The memory 160 may be configured to store computer readable/executable instructions (or firmware). The instructions, when executed by the processor 102 , may perform a number of steps.

The signal PIXELD_A-PIXELD_N may comprise raw pixel data providing a field of view captured by the lenses 150 a - 150 n . The processor 102 may be configured to generate video frames from the pixel data PIXELD_A-PIXELD_N. The video frames generated by the processor 102 may be used internal to the processor 102 (e.g., to perform video encoding, video transcoding, perform computer vision operations, etc.). In some embodiments, the video frames may be communicated to the memory 160 for temporary storage. The processor 102 may be configured to generate encoded video frames and communicate the encoded video frames to the communication device 156 as the signal ENC_VIDEO.

The processor 102 may be configured to make decisions based on analysis of the video frames generated from the signals PIXELD_A-PIXELD_N. The processor 102 may generate the signal ENC_VIDEO, the signal DATA, the signal DIR_AOUT and/or other signals (not shown). The signal ENC_VIDEO, the signal DATA and/or the signal DIR_AOUT may each be generated (in part) based on one or more decisions made and/or functions performed by the processor 102 . The decisions made and/or functions performed by the processor 102 may be determined based on data received by the <figure-callout id="102" label="processor" filenames="US11594254-20230228-D00002.png,US11594254-20230228-D00003.png" state="{{s

CLAIMS

Claims ( 20 )

The invention claimed is:

1. An apparatus comprising:

an interface configured to receive pixel data generated by a capture device; and

a processor configured to (i) receive said pixel data from said interface, (ii) process said pixel data arranged as video frames, (iii) perform computer vision operations on said video frames to generate metadata about characteristics of objects detected, (iv) determine whether said metadata corresponds to a user-defined event and (v) generate encoded video frames from a video stream of said video frames, wherein

(a) said video frames are tagged with said metadata comprising (i) an event label if said characteristics of said objects corresponds to said user-defined event and (ii) a non-event label if said characteristics of said objects do not correspond to said user-defined event,

(b) said video stream comprises at least one group of pictures comprising said event label, and

(c) said encoded video frames are generated by (i) selecting a number of said video frames from each of a plurality of groups of pictures comprising said video frames with said non-event label, and (ii) selecting more than said number of said video frames from said group of pictures comprising said video frames with said event label.

2. The apparatus according to claim 1 , wherein (i) said metadata further comprises a timestamp, (ii) each of said plurality of groups of pictures comprising said video frames with said non-event label has a first duration of said video frames and (iii) said video frames selected for said encoded video frames are arranged according to said timestamp.

3. The apparatus according to claim 1 , wherein (i) said encoded video frames are communicated to a cloud storage service, (ii) said apparatus and said capture device are implemented on an edge device and (iii) said edge device communicates with said cloud storage service.

4. The apparatus according to claim 1 , wherein (i) all of said video frames from said group of pictures comprising said video frames with said event label are selected and (ii) a frame rate of a portion of said encoded video frames that comprise said group of pictures comprising said video frames with said event label is the same as a frame rate of said video stream.

5. The apparatus according to claim 1 , wherein said encoded video frames provide a smart timelapse video stream comprising a video display speed that is automatically adjusted by said processor in response to content in said video frames.

6. The apparatus according to claim 1 , wherein (i) a user selects said objects that correspond to said user-defined event using a smartphone app and (ii) said user-defined event comprises a classification of said objects using said computer vision operations as at least one a person, a license plate, an animal depending on a selection by said user on said smartphone app.

7. The apparatus according to claim 1 , wherein a first sample of said video frames in said encoded video frames comprises a condensed version of said video frames with a high display speed for content in said plurality of groups of pictures comprising said video frames with said non-event label for said user-defined event.

8. The apparatus according to claim 7 , wherein (i) a second sample of said video frames in said encoded video frames comprises said video frames with a normal display speed for content in said group of pictures comprising said video frames with said event label and (ii) a third sample of said video frames in said encoded video frames comprises said condensed version of said video frames with said high display speed for content in said plurality of groups of pictures comprising said non-event label.

9. The apparatus according to claim 1 , wherein said event label is added to said metadata for said video frames by a convolutional neural network module implemented using a plurality of hardware engines in said processor.

10. The apparatus according to claim 1 , wherein (i) said computer vision operations are configured to detect said objects by performing feature extraction based on neural network weight values for each of a plurality of visual features that are associated with said objects extracted from said video frames and (ii) said neural network weight values are determined in response to an analysis of training data by said processor prior to said feature extraction.

11. The apparatus according to claim 1 , wherein said processor is further configured to (i) detect a first occurrence of said user-defined event in said video frames and a second occurrence of said user-defined event in said video frames, (ii) generate a first file of said encoded video frames comprising selections from said plurality of groups of pictures comprising said video frames with said non-event label comprising a timestamp before said first occurrence of said user-defined event and selections of said group of pictures comprising said event label corresponding to said first occurrence of said user-defined event and (iii) generate a second file of said encoded video frames comprising selections from said plurality of groups of pictures comprising said video frames with said non-event label comprising said timestamp after said first occurrence of said user-defined event and selections of said group of pictures comprising said event label corresponding to said second occurrence of said user-defined event.

12. A method for generating an encoded video stream, comprising the steps of:

receiving pixel data generated by a capture device,

generating video frames in response to said pixel data,

performing computer vision operations on said video frames to detect objects,

generating metadata for each of said video frames in response to performing a classification of said objects detected based on characteristics of said objects,

determining whether said classification of said objects corresponds to a user-defined event; and

generating encoded video frames from a video stream of said video frames, wherein

(a) said encoded video frames are communicated to a cloud storage service,

(b) said video frames are tagged with said metadata comprising (i) a label for said user-defined event if said classification of said objects corresponds to said user-defined event and (ii) a timestamp,

(c) said video stream comprises a plurality of groups of pictures comprising said video frames without said label for said user-defined event, each having a first duration of said video frames,

(d) said video stream comprises at least one group of pictures comprising said video frames with said label for said user-defined event,

(e) said encoded video frames are generated by (i) selecting one video frame from each of said plurality of groups of pictures comprising said video frames without said label for said user-defined event, and (ii) selecting more than one of said video frames from said group of pictures comprising said video frames with said label for said user-defined event, and

(f) said video frames selected for said encoded video frames are arranged according to said timestamp.

13. The method according to claim 12 , wherein (i) all of said video frames from said group of pictures comprising said video frames with said label for said user-defined event are selected and (ii) a frame rate of a portion of said encoded video frames that comprise said group of pictures comprising said video frames with said label for said user-defined event is the same as a frame rate of said video stream.

14. The method according to claim 12 , wherein a frame rate of a portion of said encoded video frames that comprise said group of pictures comprising said video frames with said label for said user-defined event is the same as a frame rate of said video stream.

15. The method according to claim 12 , wherein said video frames selected comprise any of an I-frame, a B-frame, or a P-frame.

16. The method according to claim 12 , further comprising the step of accessing said encoded video frames stored on said cloud storage service using a smartphone app.

17. The method according to claim 16 , further comprising the steps of:

storing a viewed timestamp corresponding to a last time that a user viewed said encoded video frames; and

presenting said encoded video frames stored by said cloud storage service to said user since said last time that said user viewed said encoded video frames.

18. The method according to claim 12 , wherein (i) a user selects said classification of said objects that correspond to said user-defined event using a smartphone app and (ii) said user-defined event comprises said classification of said objects using said computer vision operations as at least one a person, a license plate, an animal depending on a selection by said user on said smartphone app.

19. The method according to claim 12 , wherein a first sample of said video frames in said encoded video frames comprises a condensed version of said video frames with high display speed for content in said plurality of groups of pictures comprising said video frames without said label for said user-defined event.

20. The method according to claim 19 , wherein (i) a second sample of said video frames in said encoded video frames comprises said video frames with a normal display speed for content in said group of pictures comprising said video frames with said label for said user-defined event and (ii) a third sample of said video frames in said encoded video frames comprises said condensed version of said video frames with said high display speed for content in said plurality of groups of pictures comprising said video frames without said label for said user-defined event.

US17/717,060

2020-08-19

2022-04-09

Event/object-of-interest centric timelapse video generation on camera device with the assistance of neural network input

Active

US11594254B2

( en )

Priority Applications (1)

Application Number

Priority Date

Filing Date

Title

US17/717,060

US11594254B2

( en )

2020-08-19

2022-04-09

Event/object-of-interest centric timelapse video generation on camera device with the assistance of neural network input

Applications Claiming Priority (4)

Application Number

Priority Date

Filing Date

Title

CN202010836914.7A

CN114079820B

( en )

2020-08-19

2020-08-19

Video generation by recording events/objects of interest-centered intervals entered on a camera device by means of a neural network

CN202010836914.7

2020-08-19

US17/126,108

US11373685B2

( en )

2020-08-19

2020-12-18

Event/object-of-interest centric timelapse video generation on camera device with the assistance of neural network input

US17/717,060

US11594254B2

( en )

2020-08-19

2022-04-09

Event/object-of-interest centric timelapse video generation on camera device with the assistance of neural network input

Related Parent Applications (1)

Application Number

Title

Priority Date

Filing Date

US17/126,108

Continuation

US11373685B2

( en )

2020-08-19

2020-12-18

Event/object-of-interest centric timelapse video generation on camera device with the assistance of neural network input

Publications (2)

Publication Number

Publication Date

US20220230663A1

US20220230663A1 ( en )

2022-07-21

US11594254B2

true

US11594254B2 ( en )

2023-02-28

Family

ID=80269021

Family Applications (2)

Application Number

Title

Priority Date

Filing Date

US17/126,108

Active

US11373685B2

( en )

2020-08-19

2020-12-18

Event/object-of-interest centric timelapse video generation on camera device with the assistance of neural network input

US17/717,060

Active

US11594254B2

( en )

2020-08-19

2022-04-09

Event/object-of-interest centric timelapse video generation on camera device with the assistance of neural network input

Family Applications Before (1)

Application Number

Title

Priority Date

Filing Date

US17/126,108

Active

US11373685B2

( en )

2020-08-19

2020-12-18

Event/object-of-interest centric timelapse video generation on camera device with the assistance of neural network input

Country Status (2)

Country

Link

US

( 2 )

US11373685B2

( en )

CN

( 1 )

CN114079820B

( en )

Families Citing this family (21)

* Cited by examiner, † Cited by third party

Publication number

Priority date

Publication date

Assignee

Title

US20230142015A1

( en )

*

2020-11-11

2023-05-11

Milestone Systems A/S

Video surveillance system, computer-implemented video management process, and non-transitory computer readable storage medium

US20220164569A1

( en )

*

2020-11-26

2022-05-26

POSTECH Research and Business Development Foundation

Action recognition method and apparatus based on spatio-temporal self-attention

US11468676B2

( en )

*

2021-01-08

2022-10-11

University Of Central Florida Research Foundation, Inc.

Methods of real-time spatio-temporal activity detection and categorization from untrimmed video segments

US11893792B2

( en )

*

2021-03-25

2024-02-06

Adobe Inc.

Integrating video content into online product listings to demonstrate product features

EP4315856A1

( en )

*

2021-03-25

2024-02-07

Sony Semiconductor Solutions Corporation

Circuitries and methods

US12299989B2

( en )

*

2021-04-15

2025-05-13

Objectvideo Labs, Llc

Monitoring presence or absence of an object using local region matching

CN116721141B

( en )

*

2022-02-28

2026-02-10

安霸国际有限合伙企业

Adaptive facial depth map generation

CN114679607B

( en )

*

2022-03-22

2024-03-05

深圳云天励飞技术股份有限公司

Video frame rate control method and device, electronic equipment and storage medium

US20250324127A1

( en )

*

2022-06-03

2025-10-16

Clearobject Corporation

Edge device video analysis system

US20240021018A1

( en )

*

2022-07-13

2024-01-18

The Board Of Trustees Of The Leland Stanford Junior University

Systems and Methods for Recognizing Human Actions from Privacy-Preserving Optics

US12519924B2

( en )

2022-08-31

2026-01-06

Snap Inc.

Multi-perspective augmented reality experience

US12322052B2

( en )

2022-08-31

2025-06-03

Snap Inc.

Mixing and matching volumetric contents for new augmented reality experiences

US12449891B2

( en )

2022-08-31

2025-10-21

Snap Inc.

Timelapse re-experiencing system

US12267482B2

( en )

2022-08-31

2025-04-01

Snap Inc.

Controlling and editing presentation of volumetric content

US12282604B2

( en )

2022-08-31

2025-04-22

Snap Inc.

Touch-based augmented reality experience

US12417593B2

( en )

2022-08-31

2025-09-16

Snap Inc.

Generating immersive augmented reality experiences from existing images and videos

CN116389811B

( en )

*

2023-03-10

2025-07-15

东莞市九鼎实业有限公司

Synchronous control method and system for distributed video image stitching

IT202300008532A1

( en )

*

2023-05-02

2024-11-02

Vlab S R L

Method for generating a time-lapse video and associated generating device

US12556771B1

( en )

*

2023-07-11

2026-02-17

Spincast, Inc.

Video system with intra-video user input and related methods

US12294809B1

( en )

*

2024-04-05

2025-05-06

AGI7 Inc.

AI-powered cloud-native network video recorder (NVR)

WO2026009040A1

( en )

*

2024-06-30

2026-01-08

Four Drobotics Corporation

System and method for real-time artificial intelligence-based video compression and decompression

Citations (5)

* Cited by examiner, † Cited by third party

Publication number

Priority date

Publication date

Assignee

Title

US20160004390A1

( en )

*

2014-07-07

2016-01-07

Google Inc.

Method and System for Generating a Smart Time-Lapse Video Clip

US20180255273A1

( en )

*

2017-03-01

2018-09-06

Adobe Systems Incorporated

Photometric Stabilization for Time-Compressed Video

US20190130188A1

( en )

*

2017-10-26

2019-05-02

Qualcomm Incorporated

Object classification in a video analytics system

US20190364208A1

( en )

*

2019-04-17

2019-11-28

Lg Electronics Inc.

Video correction method and device

US11140292B1

( en )

*

2019-09-30

2021-10-05

Gopro, Inc.

Image capture device for generating time-lapse videos

Family Cites Families (15)

* Cited by examiner, † Cited by third party

Publication number

Priority date

Publication date

Assignee

Title

EP2407943B1

( en )

*

2010-07-16

2016-09-28

Axis AB

Method for event initiated video capturing and a video camera for capture event initiated video

US9171075B2

( en )

*

2010-12-30

2015-10-27

Pelco, Inc.

Searching recorded video

US9992443B2

( en )

*

2014-05-30

2018-06-05

Apple Inc.

System and methods for time lapse video acquisition and compression

EP3239946A1

( en )

*

2015-03-16

2017-11-01

Axis AB

Method and system for generating an event video se-quence, and camera comprising such system

WO2016192079A1

( en )

*

2015-06-04

2016-12-08

Intel Corporation

Adaptive batch encoding for slow motion video recording

CN108351965B

( en )

*

2015-09-14

2022-08-02

罗技欧洲公司

User interface for video summary

US20170300751A1

( en )

*

2016-04-19

2017-10-19

Lighthouse Al, Inc.

Smart history for computer-vision based security system

US10380429B2

( en )

*

2016-07-11

2019-08-13

Google Llc

Methods and systems for person detection in a video feed

KR102560308B1

( en )

*

2016-12-05

2023-07-27

모토로라 솔루션즈, 인크.

System and method for exterior search

EP3343561B1

( en )

*

2016-12-29

2020-06-24

Axis AB

Method and system for playing back recorded video

WO2018174505A1

( en )

*

2017-03-20

2018-09-27

Samsung Electronics Co., Ltd.

Methods and apparatus for generating video content

US10083360B1

( en )

*

2017-05-10

2018-09-25

Vivint, Inc.

Variable rate time-lapse with saliency

US10536700B1

( en )

*

2017-05-12

2020-01-14

Gopro, Inc.

Systems and methods for encoding videos based on visuals captured within the videos

US10375407B2

( en )

*

2018-02-05

2019-08-06

Intel Corporation

Adaptive thresholding for computer vision on low bitrate compressed video streams

CN110460880B

( en )

*

2019-08-09

2021-08-31

东北大学

Adaptive Transmission Method of Industrial Wireless Streaming Media Based on Particle Swarm and Neural Network

2020

2020-08-19

CN

CN202010836914.7A

patent/CN114079820B/en

active

Active

2020-12-18

US

US17/126,108

patent/US11373685B2/en

active

Active

2022

2022-04-09

US

US17/717,060

patent/US11594254B2/en

active

Active

Patent Citations (5)

* Cited by examiner, † Cited by third party

Publication number

Priority date

Publication date

Assignee

Title

US20160004390A1

( en )

*

2014-07-07

2016-01-07

Google Inc.

Method and System for Generating a Smart Time-Lapse Video Clip

US20180255273A1

( en )

*

2017-03-01

2018-09-06

Adobe Systems Incorporated

Photometric Stabilization for Time-Compressed Video

US20190130188A1

( en )

*

2017-10-26

2019-05-02

Qualcomm Incorporated

Object classification in a video analytics system

US20190364208A1

( en )

*

2019-04-17

2019-11-28

Lg Electronics Inc.

Video correction method and device

US11140292B1

( en )

*

2019-09-30

2021-10-05

Gopro, Inc.

Image capture device for generating time-lapse videos

Also Published As

Publication number

Publication date

CN114079820B

( en )

2024-11-01

US20220059132A1

( en )

2022-02-24

US11373685B2

( en )

2022-06-28

CN114079820A

( en )

2022-02-22

US20220230663A1

( en )

2022-07-21

Similar Documents

Publication

Publication Date

Title

US11373685B2

( en )

2022-06-28

Event/object-of-interest centric timelapse video generation on camera device with the assistance of neural network input

US11869241B2

( en )

2024-01-09

Person-of-interest centric timelapse video with AI input on home security camera to protect privacy

US12069251B2

( en )

2024-08-20

Smart timelapse video to conserve bandwidth by reducing bit rate of video on a camera device with the assistance of neural network input

US12499581B2

( en )

2025-12-16

Scalable coding of video and associated features

US11704908B1

( en )

2023-07-18

Computer vision enabled smart snooze home security cameras

US11880992B2

( en )

2024-01-23

Adaptive face depth image generation

EP4154511A1

( en )

2023-03-29

Maintaining fixed sizes for target objects in frames

US11798340B1

( en )

2023-10-24

Sensor for access control reader for anti-tailgating applications

CN113920010B

( en )

2025-07-15

Method and device for realizing super-resolution of image frames

US12400337B2

( en )

2025-08-26

Automatic exposure metering for regions of interest that tracks moving subjects using artificial intelligence

US11922697B1

( en )

2024-03-05

Dynamically adjusting activation sensor parameters on security cameras using computer vision

US10681313B1

( en )

2020-06-09

Home monitoring camera featuring intelligent personal audio assistant, smart zoom and face recognition features

US12002229B2

( en )

2024-06-04

Accelerating speckle image block matching using convolution techniques

US11651456B1

( en )

2023-05-16

Rental property monitoring solution using computer vision and audio analytics to detect parties and pets while preserving renter privacy

US11935257B2

( en )

2024-03-19

Adding an adaptive offset term using convolution techniques to a local adaptive binarization expression

US12069235B2

( en )

2024-08-20

Quick RGB-IR calibration verification for a mass production process

US11924555B2

( en )

2024-03-05

Intelligent auto-exposure control for RGB-IR sensor

US12452540B1

( en )

2025-10-21

Object-based auto exposure using neural network models

US20230206476A1

( en )

2023-06-29

Accelerated alignment of high-resolution image and depth map for low-bit-width floating-point representation

US11812007B2

( en )

2023-11-07

Disparity map building using guide node

US10536700B1

( en )

2020-01-14

Systems and methods for encoding videos based on visuals captured within the videos

US12610141B2

( en )

2026-04-21

Electronic image stabilization for large zoom ratio lens

US20250280198A1

( en )

2025-09-04

Large zoom ratio lens calibration for electronic image stabilization

Legal Events

Date

Code

Title

Description

2022-04-09

FEPP

Fee payment procedure

Free format text : ENTITY STATUS SET TO UNDISCOUNTED (ORIGINAL EVENT CODE: BIG.); ENTITY STATUS OF PATENT OWNER: LARGE ENTITY

2022-04-14

STPP

Information on status: patent application and granting procedure in general

Free format text : DOCKETED NEW CASE - READY FOR EXAMINATION

2022-10-04

STPP

Information on status: patent application and granting procedure in general

Free format text : NON FINAL ACTION MAILED

2022-10-23

STPP

Information on status: patent application and granting procedure in general

Free format text : RESPONSE TO NON-FINAL OFFICE ACTION ENTERED AND FORWARDED TO EXAMINER

2022-11-18

STPP

Information on status: patent application and granting procedure in general

Free format text : NOTICE OF ALLOWANCE MAILED -- APPLICATION RECEIVED IN OFFICE OF PUBLICATIONS

2023-02-08

STCF

Information on status: patent grant

Free format text : PATENTED CASE

Related documents

Record · ID 607539
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.