Voice-Based Video Tagging
US-2016055885-A1 · Feb 25, 2016 · US
US9685194B2 · US · B2
| Field | Value |
|---|---|
| Publication number | US-9685194-B2 |
| Application number | US-201414530245-A |
| Country | US |
| Kind code | B2 |
| Filing date | Oct 31, 2014 |
| Priority date | Jul 23, 2014 |
| Publication date | Jun 20, 2017 |
| Grant date | Jun 20, 2017 |
A practical reading order for non-experts. Skip the full description unless you need deep technical detail.
What the patent document calls the invention.
A short plain-language summary of the technical disclosure.
Who owns or filed the patent and who is credited as inventor.
Filing, priority, publication, and grant dates set the timeline.
The legal scope of protection — read this for what is actually claimed.
Technology tags used to group this patent with similar filings.
Prior art links and similar publications in this corpus.
Official abstract text for this publication.
Video and corresponding metadata is accessed. Events of interest within the video are identified based on the corresponding metadata, and best scenes are identified based on the identified events of interest. A video summary can be generated including one or more of the identified best scenes. The video summary can be generated using a video summary template with slots corresponding to video clips selected from among sets of candidate video clips. Best scenes can also be identified by receiving an indication of an event of interest within video from a user during the capture of the video. Metadata patterns representing activities identified within video clips can be identified within other videos, which can subsequently be associated with the identified activities.
Opening claim text (preview).
What is claimed is: 1. A method for identifying events of interest in a captured video, the method comprising: storing multiple stored speech patterns for multiple input types, the multiple stored speech patterns corresponding to a command for identifying the events of interest within the captured video, wherein the multiple stored speech patterns include a first stored speech pattern for a first input type, wherein storing the first stored speech pattern comprises: receiving, from a user, an input configuring a camera into a training mode to learn the first stored speech pattern; capturing the first stored speech pattern from the user; and storing the first stored speech pattern, wherein the first stored speech pattern is stored in response to capturing the first stored speech pattern from the user a threshold number of times; accessing a captured speech pattern, the captured speech pattern captured from the user during capture of the captured video; determining that the captured speech pattern corresponds to the first stored speech pattern; and in response to determining that the captured speech pattern corresponds to the first stored speech pattern, storing event of interest information in metadata associated with the captured video, the event of interest information identifying (i) the first input type for a first event of interest, and (ii) an event moment during the capture of the captured video at which the captured speech pattern was captured from the user. 2. The method of claim 1 , further comprising: identifying a portion of the captured video as a video clip associated with the first event of interest based on the event of interest information, the video clip comprising a first amount of the captured video occurring before the event moment and a second amount of the captured video occurring after the event moment, the first amount and the second amount being determined based on the first input type; and storing clip information indicating the association of the video clip with the first event of interest and the portion of the captured video included in the video clip. 3. The method of claim 2 , further comprising: receiving a request to generate a video summary; and generating the video summary in response to the request, the video summary comprising the video clip associated with the first event of interest. 4. The method of claim 1 , wherein the first event of interest corresponds to an activity type, wherein the first stored speech pattern identifies the activity type, and wherein storing the event of interest information comprises storing an indication of the activity type in the metadata. 5. The method of claim 1 , wherein the first stored speech pattern is specific to the user. 6. A system for identifying events of interest in a captured video, the system comprising: a processor configured by instructions to: store multiple stored speech patterns for multiple input types, the multiple stored speech patterns corresponding to a command for identifying the events of interest within the captured video, wherein the multiple stored speech patterns include a first stored speech pattern for a first input type, wherein storing the first stored speech pattern comprises: receiving, from a user, an input configuring a camera into a training mode to learn the first stored speech pattern; capturing the first stored speech pattern from the user; and storing the first stored speech pattern, wherein the first stored speech pattern is stored in response to capturing the first stored speech pattern from the user a threshold number of times; access a captured speech pattern, the captured speech pattern captured from the user during capture of the captured video; determine that the captured speech pattern corresponds to the first stored speech pattern; and in response to determining that the captured speech pattern corresponds to the first stored speech pattern, store event of interest information in metadata associated with the captured video, the event of interest information identifying (i) the first input type for a first event of interest, and (ii) an event moment during the capture of the captured video at which the captured speech pattern was captured from the user. 7. The system of claim 6 , wherein the processor is further configured to: identify a portion of the captured video as a video clip associated with the first event of interest based on the event of interest information, the video clip comprising a first amount of the captured video occurring before the event moment and a second amount of the captured video occurring after the event moment, the first amount and the second amount being determined based on the first input type; and store clip information indicating the association of the video clip with the first event of interest and the portion of the captured video included in the video clip. 8. The system of claim 7 , wherein the processor is further configured to: receive a request to generate a video summary; and generate the video summary in response to the request, the video summary comprising the video clip associated with the first event of interest. 9. The system of claim 6 , wherein the first event of interest corresponds to an activity type, wherein the first stored speech pattern identifies the activity type, and wherein storing the event of interest information comprises storing an indication of the activity type in the metadata. 10. The system of claim 6 , wherein the first stored speech pattern is specific to the user. 11. A non-transitory computer-readable storage medium storing instructions for identifying events of interest in a captured video, the instructions, when executed, causing a processor to: store multiple stored speech patterns for multiple input types, the multiple stored speech patterns corresponding to a command for identifying the events of interest within the captured video, wherein the multiple stored speech patterns include a first stored speech pattern for a first input type, wherein storing the first stored speech pattern comprises: receiving, from a user, an input configuring a camera into a training mode to learn the first stored speech pattern; capturing the first stored speech pattern from the user; and storing the first stored speech pattern, wherein the first stored speech pattern is stored in response to capturing the first stored speech pattern from the user a threshold number of times; access a captured speech pattern, the captured speech pattern captured from the user during capture of the captured video; determine that the captured speech pattern corresponds to the first stored speech pattern; and in response to determining that the captured speech pattern corresponds to the first stored speech pattern, store event of interest information in metadata associated with the captured video, the event of interest information identifying (i) the first input type for a first event of interest, and (ii) an event moment during the capture of the captured video at which the captured speech pattern was captured from the user. 12. The computer-readable storage medium of claim 11 , wherein the instructions, when executed, further causes the processor to: identify a portion of the captured video as a video clip associated with the first event of interest based on the event of interest information, the video clip comprising a first amount of the captured video occurring before the event moment and a second amount of the captured video occurring after the event moment, the first amount and the second amount being determined based on the first input type; and store clip information indicating
involving the multiplexing of an additional signal and the colour video signal · CPC title
by using information signals recorded by the same method as the main recording {(G11B27/22 takes precedence)} · CPC title
for retrieval · CPC title
Television signal processing therefor · CPC title
Indexing; Addressing; Timing or synchronising; Measuring tape travel · CPC title
Related publications grouped by family.
Answers are generated from the same data shown on this page.