re-compute and return the results according to the recognition
constraints provided in the GET-RESULT request.
The GET-RESULT request could specify constraints like a different
confidence-threshold, or n-best-list-length. This feature is
optional and the automatic speech recognition (ASR) engine may return
a status of unsupported feature.
Example:
C->S:GET-RESULT 543257 MRCP/1.0
Confidence-Threshold:90
S->C:MRCP/1.0 543257 200 COMPLETE
Content-Type:application/x-nlsml
Content-Length:276
<?xml version="1.0"?>
<result x-model="http://IdentityModel"
xmlns:xf="http://www.w3.org/2000/xforms"
grammar="session:request1@form-level.store">
<interpretation>
<xf:instance name="Person">
<Person>
<Name> Andre Roy </Name>
</Person>
</xf:instance>
<input> may I speak to Andre Roy </input>
</interpretation>
</result>
8.12. START-OF-SPEECH
This is an event from the recognizer to the client indicating that it
has detected speech. This event is useful in implementing kill-on-
barge-in scenarios when the synthesizer resource is in a different
session than the recognizer resource and, hence, is not aware of an
incoming audio source. In these cases, it is up to the client to act
as a proxy and turn around and issue the BARGE-IN-OCCURRED method to
the synthesizer resource. The recognizer resource also sends a
unique proxy-sync-id in the header for this event, which is sent to
the synthesizer in the BARGE-IN-OCCURRED method to the synthesizer.
This event should be generated irrespective of whether the
synthesizer and recognizer are in the same media server or not.
8.13. RECOGNITION-START-TIMERS
This request is sent from the client to the recognition resource when
it knows that a kill-on-barge-in prompt has finished playing. This
is useful in the scenario when the recognition and synthesizer
engines are not in the same session. Here, when a kill-on-barge-in
prompt is being played, you want the RECOGNIZE request to be
simultaneously active so that it can detect and implement kill-on-
barge-in. But at the same time, you don’t want the recognizer to
start the no-input timers until the prompt is finished. The
parameter recognizer-start-timers header field in the RECOGNIZE
request will allow the client to say if the timers should be started
or not. The recognizer should not start the timers until the client
sends a RECOGNITION-START-TIMERS method to the recognizer.
8.14. RECOGNITON-COMPLETE
This is an Event from the recognizer resource to the client
indicating that the recognition completed. The recognition result is
sent in the MRCP body of the message. The request-state field MUST
be COMPLETE indicating that this is the last event with that
request-id, and that the request with that request-id is now
complete. The recognizer context still holds the results and the
audio waveform input of that recognition until the next RECOGNIZE
request is issued. A URL to the audio waveform MAY BE returned to
the client in a waveform-url header field in the RECOGNITION-COMPLETE
event. The client can use this URI to retrieve or playback the
audio.
Example:
C->S:RECOGNIZE 543257 MRCP/1.0
Confidence-Threshold:90
Content-Type:application/grammar+xml
Content-Id:request1@form-level.store
Content-Length:104
<?xml version="1.0"?>
<!-- the default grammar language is US English -->
<grammar xml:lang="en-US" version="1.0">
<!-- single language attachment to tokens -->
<rule id="yes">
<one-of>
<item xml:lang="fr-CA">oui</item>
<item xml:lang="en-US">yes</item>
</one-of>
</rule>
<!-- single language attachment to a rule expansion -->
<rule id="request">
may I speak to
<one-of xml:lang="fr-CA">
<item>Michel Tremblay</item>
<item>Andre Roy</item>
</one-of>
</rule>
</grammar>
S->C:MRCP/1.0 543257 200 IN-PROGRESS
S->C:START-OF-SPEECH 543257 IN-PROGRESS MRCP/1.0
S->C:RECOGNITION-COMPLETE 543257 COMPLETE MRCP/1.0
Completion-Cause:000 success
Waveform-URL:http://web.media.com/session123/audio.wav
Content-Type:application/x-nlsml
Content-Length:276
<?xml version="1.0"?>
<result x-model="http://IdentityModel"
xmlns:xf="http://www.w3.org/2000/xforms"
grammar="session:request1@form-level.store">
<interpretation>
<xf:instance name="Person">
<Person>
<Name> Andre Roy </Name>
</Person>
</xf:instance>
<input> may I speak to Andre Roy </input>
</interpretation>
</result>
8.15. DTMF Detection
Digits received as DTMF tones will be delivered to the automatic
speech recognition (ASR) engine in the RTP stream according to RFC
2833 [15]. The automatic speech recognizer (ASR) needs to support
RFC 2833 [15] to recognize digits. If it does not support RFC 2833
[15], it will have to process the audio stream and extract the audio
tones from it.
9. Future Study
Various sections of the recognizer could be distributed into Digital
Signal Processors (DSPs) on the Voice Browser/Gateway or IP Phones.
For instance, the gateway might perform voice activity detection to
reduce network bandwidth and CPU requirement of the automatic speech
recognition (ASR) server. Such extensions are deferred for further
study and will not be addressed in this document.
10. Security Considerations
The MRCP protocol may carry sensitive information such as account
numbers, passwords, etc. For this reason it is important that the
client have the option of secure communication with the server for
both the control messages as well as the media, though the client is
not required to use it. If all MRCP communications happens in a
trusted domain behind a firewall, this may not be necessary. If the
client or server is deployed in an insecure network, communication
happening across this insecure network needs to be protected. In
such cases, the following additional security functionality MUST be
supported on the MRCP server. MRCP servers MUST implement Transport
Layer Security (TLS) to secure the RTSP communication, i.e., the RTSP
stack SHOULD support the rtsps: URI form. MRCP servers MUST support
Secure Real-Time Transport Protocol (SRTP) as an option to send and
receive media.
11. RTSP-Based Examples
The following is an example of a typical session of speech synthesis
and recognition between a client and the server.
Opening the synthesizer. This is the first resource for this
session. The server and client agree on a single Session ID 12345678
and set of RTP/RTCP ports on both sides.
C->S:SETUP rtsp://media.server.com/media/synthesizer RTSP/1.0
CSeq:2
Transport:RTP/AVP;unicast;client_port=46456-46457
Content-Type:application/sdp
Content-Length:190
v=0
o=- 123 456 IN IP4 10.0.0.1
s=Media Server
p=+1-888-555-1212
c=IN IP4 0.0.0.0
t=0 0
m=audio 0 RTP/AVP 0 96
a=rtpmap:0 pcmu/8000
a=rtpmap:96 telephone-event/8000
a=fmtp:96 0-15
S->C:RTSP/1.0 200 OK
CSeq:2
Transport:RTP/AVP;unicast;client_port=46456-46457;
server_port=46460-46461
Session:12345678
Content-Length:190
Content-Type:application/sdp
v=0
o=- 3211724219 3211724219 IN IP4 10.3.2.88
s=Media Server
c=IN IP4 0.0.0.0
t=0 0
m=audio 46460 RTP/AVP 0 96
a=rtpmap:0 pcmu/8000
a=rtpmap:96 telephone-event/8000
a=fmtp:96 0-15
Opening a recognizer resource. Uses the existing session ID and
ports.
C->S:SETUP rtsp://media.server.com/media/recognizer RTSP/1.0
CSeq:3
Transport:RTP/AVP;unicast;client_port=46456-46457;
mode=record;ttl=127
Session:12345678
S->C:RTSP/1.0 200 OK
CSeq:3
Transport:RTP/AVP;unicast;client_port=46456-46457;
server_port=46460-46461;mode=record;ttl=127
Session:12345678
An ANNOUNCE message with the MRCP SPEAK request initiates speech.
C->S:ANNOUNCE rtsp://media.server.com/media/synthesizer RTSP/1.0
CSeq:4
Session:12345678
Content-Type:application/mrcp
Content-Length:456
SPEAK 543257 MRCP/1.0
Kill-On-Barge-In:false
Voice-gender:neutral
Voice-category:teenager
Prosody-volume:medium
Content-Type:application/synthesis+ssml
Content-Length:104
<?xml version="1.0"?>
<speak>
<paragraph>
<sentence>You have 4 new messages.</sentence>
<sentence>The first is from <say-as
type="name">Stephanie Williams</say-as> <mark
name="Stephanie"/>
and arrived at <break/>
<say-as type="time">3:45pm</say-as>.</sentence>
<sentence>The subject is <prosody
rate="-20%">ski trip</prosody></sentence>
</paragraph>
</speak>
S->C:RTSP/1.0 200 OK
CSeq:4
Session:12345678
RTP-Info:url=rtsp://media.server.com/media/synthesizer;
seq=9810092;rtptime=3450012
Content-Type:application/mrcp
Content-Length:456
MRCP/1.0 543257 200 IN-PROGRESS
The synthesizer hits the special marker in the message to be spoken
and faithfully informs the client of the event.
S->C:ANNOUNCE rtsp://media.server.com/media/synthesizer RTSP/1.0
CSeq:5
Session:12345678
Content-Type:application/mrcp
Content-Length:123
SPEECH-MARKER 543257 IN-PROGRESS MRCP/1.0
Speech-Marker:Stephanie
C->S:RTSP/1.0 200 OK
CSeq:5
The synthesizer finishes with the SPEAK request.
S->C:ANNOUNCE rtsp://media.server.com/media/synthesizer RTSP/1.0
CSeq:6
Session:12345678
Content-Type:application/mrcp
Content-Length:123
SPEAK-COMPLETE 543257 COMPLETE MRCP/1.0
C->S:RTSP/1.0 200 OK
CSeq:6
The recognizer is issued a request to listen for the customer
choices.
C->S:ANNOUNCE rtsp://media.server.com/media/recognizer RTSP/1.0
CSeq:7
Session:12345678
RECOGNIZE 543258 MRCP/1.0
Content-Type:application/grammar+xml
Content-Length:104
<?xml version="1.0"?>
<!-- the default grammar language is US English -->
<grammar xml:lang="en-US" version="1.0">
<!-- single language attachment to a rule expansion -->
<rule id="request">
Can I speak to
<one-of xml:lang="fr-CA">
<item>Michel Tremblay</item>
<item>Andre Roy</item>
</one-of>
</rule>
</grammar>
S->C:RTSP/1.0 200 OK
CSeq:7
Content-Type:application/mrcp
Content-Length:123
MRCP/1.0 543258 200 IN-PROGRESS
The client issues the next MRCP SPEAK method in an ANNOUNCE message,
asking the user the question. It is generally RECOMMENDED when
playing a prompt to the user with kill-on-barge-in and asking for
input, that the client issue the RECOGNIZE request ahead of the SPEAK
request for optimum performance and user experience. This way, it is
guaranteed that the recognizer is online before the prompt starts
playing and the user’s speech will not be truncated at the beginning
(especially for power users).
C->S:ANNOUNCE rtsp://media.server.com/media/synthesizer RTSP/1.0
CSeq:8 Session:12345678 Content-Type:application/mrcp
Content-Length:733
SPEAK 543259 MRCP/1.0
Kill-On-Barge-In:true
Content-Type:application/synthesis+ssml
Content-Length:104
<?xml version="1.0"?>
<speak>
<paragraph>
<sentence>Welcome to ABC corporation.</sentence>
<sentence>Who would you like Talk to.</sentence>
</paragraph>
</speak>
S->C:RTSP/1.0 200 OK
CSeq:8
Content-Type:application/mrcp
Content-Length:123
MRCP/1.0 543259 200 IN-PROGRESS
Since the last SPEAK request had Kill-On-Barge-In set to "true", the
message synthesizer is interrupted when the user starts speaking, and
the client is notified.
Now, since the recognition and synthesizer resources are in the same
session, they worked with each other to deliver kill-on-barge-in. If
the resources were in different sessions, it would have taken a few
more messages before the client got the SPEAK-COMPLETE event from the
synthesizer resource. Whether the synthesizer and recognizer are in
the same session or not, the recognizer MUST generate the START-OF-
SPEECH event to the client.
The client should have then blindly turned around and issued a
BARGE-IN-OCCURRED method to the synthesizer resource. The
synthesizer, if kill-on-barge-in was enabled on the current SPEAK
request, would have then interrupted it and issued SPEAK-COMPLETE
event to the client. In this example, since the synthesizer and
recognizer are in the same session, the client did not issue the
BARGE-IN-OCCURRED method to the synthesizer and assumed that kill-
on-barge-in was implemented between the two resources in the same
session and worked.
The completion-cause code differentiates if this is normal completion
or a kill-on-barge-in interruption.
S->C:ANNOUNCE rtsp://media.server.com/media/recognizer RTSP/1.0
CSeq:9
Session:12345678
Content-Type:application/mrcp
Content-Length:273
START-OF-SPEECH 543258 IN-PROGRESS MRCP/1.0
C->S:RTSP/1.0 200 OK
CSeq:9
S->C:ANNOUNCE rtsp://media.server.com/media/synthesizer RTSP/1.0
CSeq:10
Session:12345678
Content-Type:application/mrcp
Content-Length:273
SPEAK-COMPLETE 543259 COMPLETE MRCP/1.0
Completion-Cause:000 normal
C->S:RTSP/1.0 200 OK
CSeq:10
The recognition resource matched the spoken stream to a grammar and
generated results. The result of the recognition is returned by the
server as part of the RECOGNITION-COMPLETE event.
S->C:ANNOUNCE rtsp://media.server.com/media/recognizer RTSP/1.0
CSeq:11
Session:12345678
Content-Type:application/mrcp
Content-Length:733
RECOGNITION-COMPLETE 543258 COMPLETE MRCP/1.0
Completion-Cause:000 success
Waveform-URL:http://web.media.com/session123/audio.wav
Content-Type:application/x-nlsml
Content-Length:104
<?xml version="1.0"?>
<result x-model="http://IdentityModel"
xmlns:xf="http://www.w3.org/2000/xforms"
grammar="session:request1@form-level.store">
<interpretation>
<xf:instance name="Person">
<Person>
<Name> Andre Roy </Name>
</Person>
</xf:instance>
<input> may I speak to Andre Roy </input>
</interpretation>
</result>
C->S:RTSP/1.0 200 OK
CSeq:11
C->S:TEARDOWN rtsp://media.server.com/media/synthesizer RTSP/1.0
CSeq:12
Session:12345678
S->C:RTSP/1.0 200 OK
CSeq:12
We are done with the resources and are tearing them down. When the
last of the resources for this session are released, the Session-ID
and the RTP/RTCP ports are also released.
C->S:TEARDOWN rtsp://media.server.com/media/recognizer RTSP/1.0
CSeq:13
Session:12345678
S->C:RTSP/1.0 200 OK
CSeq:13
12. Informative References
[1] Fielding, R., Gettys, J., Mogul, J., Frystyk. H., Masinter, L.,
Leach, P., and T. Berners-Lee, "Hypertext transfer protocol --
HTTP/1.1", RFC 2616, June 1999.
[2] Schulzrinne, H., Rao, A., and R. Lanphier, "Real Time Streaming
Protocol (RTSP)", RFC 2326, April 1998
[3] Crocker, D. and P. Overell, "Augmented BNF for Syntax
Specifications: ABNF", RFC 4234, October 2005.
[4] Rosenberg, J., Schulzrinne, H., Camarillo, G., Johnston, A.,
Peterson, J., Sparks, R., Handley, M., and E. Schooler, "SIP:
Session Initiation Protocol", RFC 3261, June 2002.
[5] Handley, M. and V. Jacobson, "SDP: Session Description
Protocol", RFC 2327, April 1998.
[6] World Wide Web Consortium, "Voice Extensible Markup Language
(VoiceXML) Version 2.0", W3C Candidate Recommendation, March
2004.
[7] Resnick, P., "Internet Message Format", RFC 2822, April 2001.
[8] Bradner, S., "Key words for use in RFCs to Indicate Requirement
Levels", BCP 14, RFC 2119, March 1997.
[9] World Wide Web Consortium, "Speech Synthesis Markup Language
(SSML) Version 1.0", W3C Candidate Recommendation, September
2004.
[10] World Wide Web Consortium, "Natural Language Semantics Markup
Language (NLSML) for the Speech Interface Framework", W3C
Working Draft, 30 May 2001.
[11] World Wide Web Consortium, "Speech Recognition Grammar
Specification Version 1.0", W3C Candidate Recommendation, March
2004.
[12] Yergeau, F., "UTF-8, a transformation format of ISO 10646", STD
63, RFC 3629, November 2003.
[13] Freed, N. and N. Borenstein, "Multipurpose Internet Mail
Extensions (MIME) Part Two: Media Types", RFC 2046, November
1996.
[14] Levinson, E., "Content-ID and Message-ID Uniform Resource
Locators", RFC 2392, August 1998.
[15] Schulzrinne, H. and S. Petrack, "RTP Payload for DTMF Digits,
Telephony Tones and Telephony Signals", RFC 2833, May 2000.
[16] Alvestrand, H., "Tags for the Identification of Languages", BCP
47, RFC 3066, January 2001.
Appendix A. ABNF Message Definitions
ALPHA = %x41-5A / %x61-7A ; A-Z / a-z
CHAR = %x01-7F ; any 7-bit US-ASCII character,
; excluding NUL
CR = %x0D ; carriage return
CRLF = CR LF ; Internet standard newline
DIGIT = %x30-39 ; 0-9
DQUOTE = %x22 ; " (Double Quote)
HEXDIG = DIGIT / "A" / "B" / "C" / "D" / "E" / "F"
HTAB = %x09 ; horizontal tab
LF = %x0A ; linefeed
OCTET = %x00-FF ; 8 bits of data
SP = %x20 ; space
WSP = SP / HTAB ; white space
LWS = [*WSP CRLF] 1*WSP ; linear whitespace
SWS = [LWS] ; sep whitespace
UTF8-NONASCII = %xC0-DF 1UTF8-CONT
/ %xE0-EF 2UTF8-CONT
/ %xF0-F7 3UTF8-CONT
/ %xF8-Fb 4UTF8-CONT
/ %xFC-FD 5UTF8-CONT
UTF8-CONT = %x80-BF
param = *pchar
quoted-string = SWS DQUOTE *(qdtext / quoted-pair )
DQUOTE